`get_dict_ascii_tree(items, prepend="", new_line=True)` immediately reassigned
`new_line` to the horizontal box-drawing glyph ("─"), shadowing the boolean
parameter. The trailing `if not new_line: text = text[1:]` — meant to strip the
leading newline — therefore tested a non-empty string (always truthy) and never
ran. Callers passing `new_line=False` (e.g. `maigret.py` printing a user's
identity data) got a stray leading blank line.
Rename the glyph local to `h_line` so the `new_line` parameter is preserved and
its strip takes effect. The tree drawing is unchanged.
Adds a regression test asserting `new_line=False` drops the leading newline
while keeping the rest of the tree identical.
Adds quick-access links to web archives (Wayback Machine and archive.is)
in profile URL report blocks, allowing users to see historical snapshots
of profile pages when they still exist but are no longer accessible.
Closes#247
Co-authored-by: yyq1043 <yyq1043@outlook.com>
Closes#2665.
Every site results entry flows through make_site_result (checking.py:788),
which initializes results_site = {} and then unconditionally populates it
(site/username/keywords/parsing_enabled/url_main/cookies up front, checker
at the end, plus status or url_user+future on every branch). It has no
return path that yields an empty dict. check_site_for_username wraps that
result and likewise never returns a falsy entry, and maigret.py does not
construct site-result dicts by hand.
So the four downstream 'if not dictionary: continue' guards (all tagged
'# TODO: fix no site data issue') were dead code — the entries they
guarded against cannot occur. Removed:
- maigret/maigret.py:101 (extract_ids_from_results)
- maigret/report.py:151 — 'if not dictionary or dictionary.get("is_similar")'
trimmed to just the is_similar check (real logic)
- maigret/report.py:467 (extended-report builder)
- maigret/report.py:609 (generate_txt_report)
- maigret/report.py:627 — 'if not site_result or not site_result.get("status")'
trimmed to just the status check (real logic)
- maigret/report.py:687 (generate_json_report)
plus the four '# TODO: fix no site data issue' comments.
No behavior change for well-formed inputs; downstream .get() calls are
safe because entries are always populated dicts. 292 tests pass.
Co-authored-by: aznikline <aznikline@users.noreply.github.com>
Closes#2666.
Issue #2666 tracked four paths flagged `# TODO: tests`. Two of them
(`executors.py` increment_progress / stop_progress) were removed by the
#2789 refactor and no longer exist. This PR covers the two that remain:
1. maigret/checking.py:247 — the ClientSession cookie_jar forwarding.
SimpleAiohttpChecker.check() builds ClientSession(cookie_jar=self.cookie_jar
if self.cookie_jar else None), but nothing asserted the jar handed to the
checker actually reached the outgoing session. Added two tests using the
existing constructor-capture pattern (_capture_clientsession): one asserts
a configured jar is forwarded verbatim, the other asserts None is passed
when no jar is configured (the else branch).
2. maigret/maigret.py:946 — extract_ids_from_results. The pre-existing
test_extract_ids_from_results was a bare expression with no `assert` (so
it silently passed regardless of return value) AND pinned its expectation
against test_db (tests/db.json), whose site set never matched the Reddit
URL and silently dropped the ids_links result. Fixed: real assertion against
default_db, plus three branch tests (cross-site username merge, empty-site
skip, empty input).
No new dependencies: aioresponses was suggested in the issue but is not in
the project — the existing monkeypatch capture pattern fits the codebase
better and adds zero deps.
Removed both `# TODO: tests` comments from the now-covered lines.
Co-authored-by: aznikline <aznikline@users.noreply.github.com>
`extract_and_group` computed the per-error-type percentage as
`round(count / len(search_res), 2) * 100` — it rounded the *fraction* to 2
decimals (i.e. to the nearest 1%) and only then scaled by 100, so `perc`
always had whole-percent granularity and was mis-rounded.
A genuine 2.5% error rate (1 error out of 40 sites) became
`round(0.025, 2) * 100 = 3.0%` and tripped the 3% "important" threshold in
`is_important()`, firing a spurious "Too many errors of type ..." warning.
Likewise 9.5% DNS errors rounded to 10.0% and crossed the DNS override.
Scale to a percentage *before* rounding: `round(count / len(search_res) * 100, 2)`.
The existing tests use integer-landing ratios (25/100, 5/100, 3/100) that are
identical under both formulas, so none of them change.
Adds a regression test for a sub-threshold non-integer rate (2.5%) staying silent.
Convert activation HTTP calls to aiohttp coroutines, await activation before retrying, and allocate independent protocol checkers per site check so concurrent retries do not overwrite shared checker state.
is_country_tag checks for a 2-letter alpha code. The field names
'country' and 'locale' are 6-7 characters and never match, so the
direct alpha_2 lookup branch was dead code and all country values fell
through to search_fuzzy unconditionally.
Change is_country_tag(k) -> is_country_tag(v) so that 2-letter ISO
codes (e.g. 'US', 'RU') use the fast direct lookup while full country
names continue to use fuzzy search.
Fixes#2752
update_site() used `s = site` inside a for-loop, which rebinds only the
local loop variable and leaves self._sites[i] unchanged. The method
returned self silently appearing to succeed while the database list
retained the original entry. This broke --auto-disable for any site
already present in the database.
Fix with enumerate so the actual list element is replaced.
Fixes#2750
Adding PDF dependency by default for web variant as a search will result in an error without installing this. In case it is not installed it maigret web will return to the landing page and will present an error in the logs.