docs: link to socid-extractor readthedocs pages (#2927) (#2967)

This commit is contained in:
felipe
2026-08-12 17:15:58 +02:00
committed by GitHub
parent b6fe0cd0e3
commit 46c90bd2cc
8 changed files with 31 additions and 15 deletions
+2 -2
View File
@@ -202,13 +202,13 @@ Site check fixes using LLM
--------------------------
.. note::
The ``LLM/`` directory at the root of the repository contains detailed instructions for editing site checks (in Markdown format): checklist, full guide to ``checkType`` / ``data.json`` / ``urlProbe``, handling false positives, searching for public JSON APIs, and the proposal log for ``socid_extractor``.
The ``LLM/`` directory at the root of the repository contains detailed instructions for editing site checks (in Markdown format): checklist, full guide to ``checkType`` / ``data.json`` / ``urlProbe``, handling false positives, searching for public JSON APIs, and the proposal log for `socid_extractor <https://socid-extractor.readthedocs.io/>`_.
Main files:
- `site-checks-playbook.md <https://github.com/soxoj/maigret/blob/main/LLM/site-checks-playbook.md>`_ — short checklist
- `site-checks-guide.md <https://github.com/soxoj/maigret/blob/main/LLM/site-checks-guide.md>`_ — detailed guide
- `socid_extractor_improvements.log <https://github.com/soxoj/maigret/blob/main/LLM/socid_extractor_improvements.log>`_ — template and entries for identity extractor improvements
- `socid_extractor_improvements.log <https://github.com/soxoj/maigret/blob/main/LLM/socid_extractor_improvements.log>`_ — template and entries for identity extractor improvements; the upstream how-to is `Adding a scheme <https://socid-extractor.readthedocs.io/en/latest/adding-a-scheme.html>`_
These files should be kept up-to-date whenever changes are made to the check logic in the code or in ``data.json``.
+1 -1
View File
@@ -50,7 +50,7 @@ installing anything locally.
Personal info gathering
-----------------------
Maigret does the `parsing of accounts webpages and extraction <https://github.com/soxoj/socid-extractor>`_ of personal info, links to other profiles, etc.
Maigret does the `parsing of accounts webpages and extraction <https://socid-extractor.readthedocs.io/en/latest/how-extraction-works.html>`_ of personal info, links to other profiles, etc.
Extracted info displayed as an additional result in CLI output and as tables in HTML and PDF reports.
Also, Maigret use found ids and usernames from links to start a recursive search.
+1 -1
View File
@@ -52,7 +52,7 @@ A working end-to-end search against the top 500 sites:
Key points:
- ``maigret_search`` is an ``async`` function — wrap it with ``asyncio.run(...)`` or ``await`` it from inside your own event loop.
- ``is_parsing_enabled=True`` turns on ``socid_extractor`` so ``result["ids_data"]`` is populated with profile fields (bio, linked accounts, uids, etc.).
- ``is_parsing_enabled=True`` turns on ``socid_extractor`` so ``result["ids_data"]`` is populated with profile fields (bio, linked accounts, uids, etc.). See its `Library usage <https://socid-extractor.readthedocs.io/en/latest/library-usage.html>`_ page for calling it directly, and `Maigret integration <https://socid-extractor.readthedocs.io/en/latest/maigret-integration.html>`_ for how the two projects fit together.
- Each entry in the returned dict has a ``"status"`` object with ``is_found()``, plus ``url_user``, ``http_status``, ``rank``, ``ids_data``, and more.
Filtering sites
@@ -645,11 +645,11 @@ msgid ""
"instructions for editing site checks (in Markdown format): checklist, "
"full guide to ``checkType`` / ``data.json`` / ``urlProbe``, handling "
"false positives, searching for public JSON APIs, and the proposal log for"
" ``socid_extractor``."
" `socid_extractor <https://socid-extractor.readthedocs.io/>`_."
msgstr ""
"仓库根目录下的 ``LLM/`` 目录中,以 Markdown 形式保存了编辑站点检查的详细指南:检查清单、关于 ``checkType`` / "
"``data.json`` / ``urlProbe`` 的完整说明、处理误报的方法、寻找公开 JSON API 的思路,以及面向 "
"``socid_extractor`` 的改动提案日志。"
"`socid_extractor <https://socid-extractor.readthedocs.io/>`_ 的改动提案日志。"
#: ../../source/development.rst:199 5d8b27e0fd99416a9e92869bc4ac64e5
msgid "Main files:"
@@ -675,11 +675,14 @@ msgstr ""
msgid ""
"`socid_extractor_improvements.log "
"<https://github.com/soxoj/maigret/blob/main/LLM/socid_extractor_improvements.log>`_"
" — template and entries for identity extractor improvements"
" — template and entries for identity extractor improvements; the upstream "
"how-to is `Adding a scheme "
"<https://socid-extractor.readthedocs.io/en/latest/adding-a-scheme.html>`_"
msgstr ""
"`socid_extractor_improvements.log "
"<https://github.com/soxoj/maigret/blob/main/LLM/socid_extractor_improvements.log>`_"
" —— 身份信息抽取器改进项的模板与记录"
" —— 身份信息抽取器改进项的模板与记录;上游的操作指南参见 `Adding a scheme "
"<https://socid-extractor.readthedocs.io/en/latest/adding-a-scheme.html>`_"
#: ../../source/development.rst:205 53286ba37db94831adda43dee21158c6
msgid ""
@@ -98,12 +98,14 @@ msgstr "个人信息收集"
#: ../../source/features.rst:53 2913acff61f8466e80dc03285b75c54b
msgid ""
"Maigret does the `parsing of accounts webpages and extraction "
"<https://github.com/soxoj/socid-extractor>`_ of personal info, links to "
"<https://socid-extractor.readthedocs.io/en/latest/how-extraction-works.html>`_"
" of personal info, links to "
"other profiles, etc. Extracted info displayed as an additional result in "
"CLI output and as tables in HTML and PDF reports. Also, Maigret use found"
" ids and usernames from links to start a recursive search."
msgstr ""
"Maigret 会\\ `解析账号网页并抽取 <https://github.com/soxoj/socid-extractor>`_ "
"Maigret 会\\ `解析账号网页并抽取 "
"<https://socid-extractor.readthedocs.io/en/latest/how-extraction-works.html>`_ "
"个人信息、指向其它主页的链接等内容。抽取结果会以附加信息的形式出现在 CLI 输出中,并在 HTML 和 PDF "
"报告中以表格呈现。此外,Maigret 还会用从链接中发现的 ID 和用户名,启动递归搜索。"
@@ -66,10 +66,18 @@ msgstr ""
msgid ""
"``is_parsing_enabled=True`` turns on ``socid_extractor`` so "
"``result[\"ids_data\"]`` is populated with profile fields (bio, linked "
"accounts, uids, etc.)."
"accounts, uids, etc.). See its `Library usage "
"<https://socid-extractor.readthedocs.io/en/latest/library-usage.html>`_ "
"page for calling it directly, and `Maigret integration "
"<https://socid-extractor.readthedocs.io/en/latest/maigret-integration.html>`_"
" for how the two projects fit together."
msgstr ""
"``is_parsing_enabled=True`` 会启用 ``socid_extractor``,从而把个人主页中"
"的字段(个人简介、关联账号、各类 uid 等)填充到 ``result[\"ids_data\"]``\ 。"
"直接调用该库的方法参见 `Library usage "
"<https://socid-extractor.readthedocs.io/en/latest/library-usage.html>`_\\ ,"
"两个项目如何协同参见 `Maigret integration "
"<https://socid-extractor.readthedocs.io/en/latest/maigret-integration.html>`_\\ 。"
#: ../../source/library-usage.rst:56 886c9269cbf346819fd029eb0c86b1ff
msgid ""
@@ -87,12 +87,14 @@ msgstr "由此衍生出两个想法:"
#: ../../source/philosophy.rst:35 bb7a5044518943068fff752a527be1eb
msgid ""
"`socid-extractor <https://github.com/soxoj/socid-extractor>`_ — a library"
" focused on pulling structured identity data (user IDs, full names, "
"`socid-extractor <https://github.com/soxoj/socid-extractor>`_ "
"(`documentation <https://socid-extractor.readthedocs.io/>`_) — a library "
"focused on pulling structured identity data (user IDs, full names, "
"linked accounts, bios, timestamps, etc.) out of account pages and public "
"API responses, so that finding an account is not the end of the pipeline."
msgstr ""
"`socid-extractor <https://github.com/soxoj/socid-extractor>`_ —— 一个专注于从账号页面和公开 API 响应中抽取结构化身份数据(用户 ID、真实姓名、关联账号、个人简介、时间戳等)的库,让“找到账号”不再是流水线的终点。"
"`socid-extractor <https://github.com/soxoj/socid-extractor>`_\\ (`文档 "
"<https://socid-extractor.readthedocs.io/>`_\\ )—— 一个专注于从账号页面和公开 API 响应中抽取结构化身份数据(用户 ID、真实姓名、关联账号、个人简介、时间戳等)的库,让“找到账号”不再是流水线的终点。"
#: ../../source/philosophy.rst:38 d8778348b35e4514998f83f1ba7bb49a
msgid ""
+2 -1
View File
@@ -32,7 +32,8 @@ For a broader landscape of username-checking tools, see the curated
Two ideas grew out of that research:
- `socid-extractor <https://github.com/soxoj/socid-extractor>`_ — a library focused on pulling
- `socid-extractor <https://github.com/soxoj/socid-extractor>`_
(`documentation <https://socid-extractor.readthedocs.io/>`_) — a library focused on pulling
structured identity data (user IDs, full names, linked accounts, bios, timestamps, etc.) out of
account pages and public API responses, so that finding an account is not the end of the pipeline.
- **Maigret** itself — which started as a fork of