The slow 'first pass' was one sequential 'gh pr view' per open PR (~one
network round-trip each), so it is I/O-bound. Fetch them concurrently with
asyncio.create_subprocess_exec, bounded by a semaphore (default 10,
override via PR_FILE_MAP_CONCURRENCY) to stay polite to the GitHub API.
Output is unchanged and deterministic (results keyed back in PR order);
errors still fail fast with the same messages.
* pr_file_map.py: add ignore_pull_request list, progress passes, dated title
- Add module-level ignore_pull_request: list[int]; get_open_prs() filters
those PR numbers out (e.g. [123, 456, 789] skips #123, #456, #789).
- Drop the redundant DIRECTORY.md comment above DIRECTORY_FILE (the module
docstring and render_directory_section already explain it).
- Fold the generation datetime into the H1 title:
'# Open Pull Request File Map: 16 Sep 2026 at 21:45 UTC'.
- Show progress on stderr via other/cheap_progress.py's progress() for the
'First pass' (fetch each PR's files) and 'Second pass' (classify files).
* Change ignore_pull_request from list to set
---------
Co-authored-by: Christian Clauss <cclauss@me.com>
2026-09-17 15:58:43 +02:00
Christian Claussandpre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* pr_file_map.py: Use a human-readable date format
Use a human-readable date format.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Move DIRECTORY.md to its own section in pr_file_map.py
DIRECTORY.md is auto-generated and touched by nearly every open PR, so it
dominated the 'possible merge conflicts' list and distracted maintainers.
Pull it out into a dedicated section at the very bottom of the report that
separates PRs whose only overlap is DIRECTORY.md (safe to 'accept both' in
the GitHub UI) from PRs that also collide on real source files.
Also render the Script path relative to the git root instead of an absolute
machine-specific path.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Add a 'files touched by more than one open PR' section (sorted with the
most-contested files first) so overlapping PRs -- the likely merge-conflict
hot spots -- are visible at a glance when deciding what to land.
Also report two distinct file totals: 'file touches' (every PR x file pair)
and 'distinct files touched'. Only the distinct total equals existing +
missing, which fixes the earlier single count that double-counted files
edited by multiple PRs.
* fix:bucket count type
* fixed TypeError
* Refine docstring and adjust return statement
Updated docstring for partition_liked_list method to clarify behavior. Changed return statement from None to None for consistency.
* Add script to map open PRs to modified files
This script lists all open pull requests in the current directory's git repository and maps each file touched by any open PR to its corresponding PR numbers. It outputs the results in GitHub-flavored Markdown format.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Add noqa comments for subprocess calls
Add noqa comments to suppress specific linting warnings.
---------
Co-authored-by: Christian Clauss <cclauss@me.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* fix(ci): count open issues reliably and add a Hacktoberfest countdown
The tracker's "Open issues" line was showing the pull-request total (e.g.
621) instead of the issue total (107). GitHub's `/search/issues`
`is:issue` / `is:pr` qualifiers are unreliable on this large, high-churn
repo -- some runs return the PR pool for *both* queries, so the two lines
printed the same number.
Count without the flaky qualifiers instead:
- open PRs = the `Link: rel="last"` page number of `/repos/{repo}/pulls`
(deterministic, page-numbered pagination);
- open issues = the repo endpoint's `open_issues_count` (issues + PRs)
minus the open-PR total -- self-checking and stable.
Also add the requested Hacktoberfest countdown to the stats block:
- days until 2026-10-01;
- issues to close per day to clear the backlog;
- PRs to merge or close per day to clear the backlog.
Per-day figures round up (finishing a day early beats a day late) and
degrade to a clear message once Hacktoberfest starts, so the block never
divides by zero on the final day.
* chore: re-trigger keeper after marking PR checklist
* fix(ci): open a PR to persist Hacktoberfest tracker instead of pushing to protected master
* ci: use bundled gh CLI instead of peter-evans/create-pull-request
Per @cclauss / zizmor 'superfluous actions' audit, persist the rolling
tracker PR with the gh CLI rather than a third-party action.
* fix(ci): make tracker refresh degrade gracefully when rate limited
The dry run was failing because a run can exhaust the GITHUB_TOKEN's
1000/hour-per-repo budget (shared across concurrent runs) — chiefly the
awaiting-reviews directory scan. A single exhausted request then raised
and killed the whole job.
- _request now honours Retry-After (secondary limits) and, once retries
are exhausted, raises BestEffortError instead of a bare RuntimeError.
- Row resolution, the directory scan, and the search counts catch
BestEffortError and degrade (keep the row / mark the stat unavailable)
instead of failing. Only the post-Oct-1 retirement exits non-zero.
- Trim the directory scan to 120 PRs and CONCURRENCY to 5 to stay well
under the shared budget in the first place.
* ci: dry-run hacktoberfest prep on push/PR, add path filters
* ci: use double quotes in path filters (prettier)
* fix: resolve tracker rows via /issues so issue rows don't 404
The tracker's 'Open issues' section lists issue numbers; querying them
against /pulls/{n} returns 404 and crashed the whole refresh. Query the
unified /issues/{n} endpoint instead, which resolves for both PRs and
issues; a row is 'merged' only when it's a PR with merged_at set.
The daily prep job spent ~16 minutes because it looked up the changed
files of every open 'awaiting reviews' PR one request at a time. Those
lookups are independent, so fire them concurrently through a single
httpx2.AsyncClient bounded by a small semaphore, and likewise resolve the
tracked-row PR states and the issue/PR counts concurrently. Report
progress to stderr (flushed) so a human watching the Actions log can see
the job is alive. Runtime drops from ~16 min of serial round-trips to
well under a minute.
* ci: daily Hacktoberfest 2026 prep tracker refresh (cron 11:50 UTC)
Adds .github/workflows/hacktoberfest_prep.yml (schedule: 50 11 * * *) and
scripts/hacktoberfest_prep_update.py. Each run:
- ticks tracked PR rows in docs/hacktober_2026_prep.md that are now
merged/closed (`[ ]` -> `[x]`),
- rewrites an 'Automated statistics' section with the current open issue and
open PR counts plus the top three directories with the most open
'awaiting reviews' PRs,
- exits non-zero once Hacktoberfest 2026 has begun (>= 2026-10-01), so the
prep window closing is loud and the job gets retired.
Standard library only; uses the Actions GITHUB_TOKEN. Refs #15081.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* style: wrap implicit string concatenations (ISC004)
* refactor: use httpx2 for API calls, drop unneeded future import
Address review feedback on the Hacktoberfest prep cron:
- Switch the tracker script from urllib to httpx2, the repo's standard
HTTP client, and add an install step to the workflow.
- Drop 'from __future__ import annotations' (unnecessary on Python >= 3.14t).
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Add a table of contents to DIRECTORY.md
Generate a linked table of contents of the top-level sections at the top
of DIRECTORY.md so readers can jump straight to a category, as suggested
in #13239 / #13111. The list is built in scripts/build_directory_md.py
(with a new md_anchor helper + doctests) so DIRECTORY.md stays fully
auto-generated.
* Address review: numbered ToC + link section headings to their directories
- Table of Contents is now a numbered list, so the final number shows the
total count of algorithm folders at a glance.
- Each top-level section heading links to its algorithm directory (e.g.
## [Sorts](sorts)), so clicking a section title jumps straight to the folder.
* deps: migrate from httpx to httpx2 (pydantic's maintained fork)
Mechanical rename of httpx -> httpx2 (API-compatible fork of httpx 0.28.1):
pyproject.toml deps, PEP 723 inline-script headers, and all import/call sites.
Excludes uv.lock (the keeper's allow-list rejects .lock files); the lock
refresh needs a separate maintainer-merged PR.
Refs #15081
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* deps: drop tweepy + migrate remaining requests refs to httpx2
- maths/allocation_number.py: docstring example uses httpx2, not requests
- web_programming/get_imdbtop.py.DISABLED: import httpx2 instead of requests
- remove web_programming/get_user_tweets.py.DISABLED (a Twitter API how-to,
not an algorithm) and drop the tweepy dependency that was its only user and
the last high-level dep pulling in requests
- uv.lock intentionally untouched (keeper allow-list)
Refs #15081
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* scripts: de-duplicate and harden the close_pull_requests_with_*.sh backlog jobs
Extract the five byte-for-byte-identical backlog-closing scripts into one
parameterized close_pull_requests_with_label.sh and make each named script a
thin wrapper. Addresses the recommendations from #15168:
- Correctness: filter by label server-side (gh pr list --label) instead of
listing all ~900 open PRs and matching client-side, so no PR is skipped by
an arbitrary --limit cap.
- De-duplication: one implementation removes the drift between copies (some
had sleep 2, one had it commented out, two had none).
- Throttling: a single, deliberate SLEEP (default 2s) between closes.
- Safety rails: set -euo pipefail, explicit --repo, and a DRY_RUN=1 preview
mode that prints what would close before a maintainer commits.
- Machine-readable summary (CLOSED_COUNT/CLOSED_PRS) so the Hacktoberfest
tracker can be updated from script output.
Follow-up to #15081.
* updating DIRECTORY.md
---------
Co-authored-by: priya-sundaram-dev <oc-409d01@agentmail.to>
Co-authored-by: cclauss <cclauss@users.noreply.github.com>
* [pre-commit.ci] pre-commit autoupdate
updates:
- [github.com/astral-sh/ruff-pre-commit: v0.8.6 → v0.9.1](https://github.com/astral-sh/ruff-pre-commit/compare/v0.8.6...v0.9.1)
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Update maths/dual_number_automatic_differentiation.py
* Update maths/dual_number_automatic_differentiation.py
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Update dual_number_automatic_differentiation.py
* Update dual_number_automatic_differentiation.py
* No <fin-streamer> tag with the specified data-test attribute found.
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Christian Clauss <cclauss@me.com>
* Enable ruff S113 rule
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Enable ruff PGH003 rule
* Fix
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Fix
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Enable ruff SIM102 rule
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Fix
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
* Added the matrix_exponentiation.py file in maths directory
* Implemented the requested changes
* Update matrix_exponentiation.py
* resolve merge conflict with upstream branch
* add new line at end of file
* add wavelet_tree
* fix isort issue
* updating DIRECTORY.md
* fix variable names in wavelet_tree and correct typo
* Add type hints and variable renaming
* Update data_structures/binary_tree/wavelet_tree.py
Add doctests to placate the algorithm-bot, thanks to @cclauss.
Co-authored-by: Christian Clauss <cclauss@me.com>
* Move doctest to individual functions and reformat code
* Move common test array to the global scope and reuse in tests
* MMove test array to global scope and minor linting changes
* Correct the failing pytest tests
* MUse built-in list for type annotation
* Update wavelet_tree.py
* types-requests
* updating DIRECTORY.md
* Update wavelet_tree.py
* # type: ignore
* # type: ignore
* Update decrypt_caesar_with_chi_squared.py
* ,
* Update decrypt_caesar_with_chi_squared.py
Co-authored-by: Christian Clauss <cclauss@me.com>
Co-authored-by: github-actions <${GITHUB_ACTOR}@users.noreply.github.com>
Co-authored-by: Aniruddha Bhattacharjee <aniruddha@Aniruddhas-MacBook-Air.local>
* Update validate solution script to fetch only submitted solution
* Update workflow file with the updated PE script
* Fix: do not fetch `validate_solutions.py` script
* Update script to use the requests package for API calls
* Fix: install requests module
* Pytest ignore scripts/ directory