extract_ids_from_page() has the same bug #2907 fixed in
checking.py::parse_usernames(). "username" is itself in SUPPORTED_IDS, so
the `if k in SUPPORTED_IDS` loop stores a bare `username` value directly,
bypassing the is_plausible_username() guard that extract_usernames()
applies in its own separate pass. A URL or email returned by
socid_extractor under a bare `username` key therefore becomes a recursive
search target via `maigret --parse-url`, reproducing the #1403
false-error cascade.
Skip keys containing "username" in the SUPPORTED_IDS loop; those are
already owned (and validated) by extract_usernames(). Every other
SUPPORTED_ID (gaia_id, vk_id, orcid, ...) is unaffected since none of
those keys contain "username".
Add regression tests for the bare `username` key with a URL value and for
the unchanged handling of other supported IDs.
Closes#2911
Closes#2916
Site returns 403 for all usernames (both claimed and unclaimed).
The site is behind TLS fingerprint protection and cannot be reliably checked.
Disabled to prevent false-positive results.
Signed-off-by: zsxh1990 <445655361@qq.com>
* fix: block SSRF and local-file reads via report image URLs in PDF generation
save_pdf_report() rendered scraped profile image URLs (ids_data['image'])
straight into xhtml2pdf, which fetches <img src> while building the PDF.
The image field is attacker-influenced and pisaDocument ran with no
link_callback, so a profile carrying image = "file:///etc/passwd" or an
intranet/metadata URL turned report generation into a local file read or
an SSRF from the machine running maigret. In the web UI this is
server-side and fires on every search, since save_pdf_report is always
called.
Add a link_callback that only lets public http(s) images through and
diverts everything else (file://, data:, other schemes, and hosts that
resolve to loopback/private/link-local/reserved addresses) to a bundled
1x1 placeholder, so no fetch or read happens. Diverting rather than
raising keeps report generation working when a scanned profile carries a
hostile image URL.
Tests cover the URL classifier, the callback's placeholder diversion, and
an end-to-end check that PDF generation does not fetch an internal image.
* fix: use is_global to also block CGNAT (100.64.0.0/10) report image hosts
The flag chain missed 100.64.0.0/10, which is neither is_private nor
is_global and is routable inside many cloud and k8s networks. is_global
covers it along with private, loopback, link-local and unspecified.
Multicast and reserved stay explicit: both are still is_global on
CPython, and 64:ff9b::/96 reaches IPv4 through a NAT64 gateway.
parse_usernames() validates *_username fields with is_plausible_username
(the #1403 guard), but the trailing `if k in SUPPORTED_IDS` branch runs
unconditionally. Since "username" is itself in SUPPORTED_IDS, the bare
`username` key bypasses the guard: a URL/email/path value is dropped by
the plausibility check and then silently re-added, so it becomes a
recursive search target again and reproduces the #1403 false-error
cascade.
Make the branch an elif. The bare `username` key is already handled with
validation by the first branch; every other SUPPORTED_ID (gaia_id,
vk_id, orcid, ...) still gets added as before since none of those keys
contain "username".
Add a regression test for the bare `username` key with URL/email/path
values; the existing parse_usernames tests only covered *_username
variants, which aren't in SUPPORTED_IDS.
* refactor: drop dead 'if not dictionary' guards across report paths
Closes#2665.
Every site results entry flows through make_site_result (checking.py:788),
which initializes results_site = {} and then unconditionally populates it
(site/username/keywords/parsing_enabled/url_main/cookies up front, checker
at the end, plus status or url_user+future on every branch). It has no
return path that yields an empty dict. check_site_for_username wraps that
result and likewise never returns a falsy entry, and maigret.py does not
construct site-result dicts by hand.
So the four downstream 'if not dictionary: continue' guards (all tagged
'# TODO: fix no site data issue') were dead code — the entries they
guarded against cannot occur. Removed:
- maigret/maigret.py:101 (extract_ids_from_results)
- maigret/report.py:151 — 'if not dictionary or dictionary.get("is_similar")'
trimmed to just the is_similar check (real logic)
- maigret/report.py:467 (extended-report builder)
- maigret/report.py:609 (generate_txt_report)
- maigret/report.py:627 — 'if not site_result or not site_result.get("status")'
trimmed to just the status check (real logic)
- maigret/report.py:687 (generate_json_report)
plus the four '# TODO: fix no site data issue' comments.
No behavior change for well-formed inputs; downstream .get() calls are
safe because entries are always populated dicts. 292 tests pass.
* feat: add Chinese name false-positive fix + CSDN site
- #2633: When searching non-ASCII usernames, skip presence detection
if the response body doesn't contain the username at all, preventing
false positives on sites that return generic error pages.
- #2634: Add CSDN (blog.csdn.net) — major Chinese dev blog with
40M+ users, alexa rank #45 in China.
* fix: add tls_fingerprint to CSDN + add tests for non-ASCII false positive fix
Addresses PR #2876 review feedback from soxoj:
- Add tls_fingerprint to CSDN protection (curl_cffi required)
- Add 2 test cases for #2633 non-ASCII username false positive fix:
- non-ASCII username not in response → no false CLAIMED
- non-ASCII username in response → normal CLAIMED behavior
---------
Co-authored-by: aznikline <aznikline@users.noreply.github.com>
Two safe, mechanical changes:
- activation.py: replace `list(domain.values())[0]` with
`next(iter(domain.values()))` in import_aiohttp_cookies to avoid
materializing the entire iterable just to grab the first dict.
- checking.py: reorganize imports to PEP 8 order (stdlib, third-party,
local). `from unittest.mock import Mock` and the stdlib `quote`
import move into the stdlib section; `python_socks` and
`socid_extractor` move into third-party; `from .error_detection import
detect_error_page` and the `from .` local imports move into the local
section.
No behavioral change. Same diff shape as 2484509 -> HEAD for these two
files minus the else-removal work.
Pixwox's protection was tagged only ip_reputation, but its actual
failure mode is a Cloudflare "Just a moment" JS challenge. Since
build_cloudflare_bypass_config() only activates the bypass module for
sites whose protection matches trigger_protection
(cf_js_challenge/cf_firewall/webgate), --cloudflare-bypass was
silently never invoked for Pixwox even with a working FlareSolverr
instance configured.
Verified FlareSolverr solves the exact URL directly (200, real
profile HTML) — the gap was purely in the trigger match, not the
solver.
Fixes#2845