Files
Python/web_programming/reddit.py
priya-sundaram-devandpre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> bfa655ecea deps: migrate from httpx to httpx2 (pydantic's maintained fork) (#15192)
* deps: migrate from httpx to httpx2 (pydantic's maintained fork)

Mechanical rename of httpx -> httpx2 (API-compatible fork of httpx 0.28.1):
pyproject.toml deps, PEP 723 inline-script headers, and all import/call sites.
Excludes uv.lock (the keeper's allow-list rejects .lock files); the lock
refresh needs a separate maintainer-merged PR.

Refs #15081

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* deps: drop tweepy + migrate remaining requests refs to httpx2

- maths/allocation_number.py: docstring example uses httpx2, not requests
- web_programming/get_imdbtop.py.DISABLED: import httpx2 instead of requests
- remove web_programming/get_user_tweets.py.DISABLED (a Twitter API how-to,
  not an algorithm) and drop the tweepy dependency that was its only user and
  the last high-level dep pulling in requests
- uv.lock intentionally untouched (keeper allow-list)

Refs #15081

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-09-05 09:26:45 +02:00

66 lines
2.2 KiB
Python

# /// script
# requires-python = ">=3.13"
# dependencies = [
# "httpx2",
# ]
# ///
from typing import Literal
import httpx2
valid_terms = set(
"""approved_at_utc approved_by author_flair_background_color
author_flair_css_class author_flair_richtext author_flair_template_id author_fullname
author_premium can_mod_post category clicked content_categories created_utc downs
edited gilded gildings hidden hide_score is_created_from_ads_ui is_meta
is_original_content is_reddit_media_domain is_video link_flair_css_class
link_flair_richtext link_flair_text link_flair_text_color media_embed mod_reason_title
name permalink pwls quarantine saved score secure_media secure_media_embed selftext
subreddit subreddit_name_prefixed subreddit_type thumbnail title top_awarded_type
total_awards_received ups upvote_ratio url user_reports""".split()
)
def get_subreddit_data(
subreddit: str,
limit: int = 1,
age: Literal["new", "top", "hot"] = "new",
wanted_data: list | None = None,
) -> dict:
"""
subreddit : Subreddit to query
limit : Number of posts to fetch
age : ["new", "top", "hot"]
wanted_data : Get only the required data in the list
"""
wanted_data = wanted_data or []
if invalid_search_terms := ", ".join(sorted(set(wanted_data) - valid_terms)):
msg = f"Invalid search term: {invalid_search_terms}"
raise ValueError(msg)
# raise_for_status() already raises httpx2.HTTPStatusError for any 4xx/5xx
# response (including 429), so no extra status check is needed here.
data = (
httpx2.get(
f"https://www.reddit.com/r/{subreddit}/{age}.json?limit={limit}",
headers={"User-agent": "A random string"},
timeout=10,
)
.raise_for_status()
.json()
)
if not wanted_data:
return {id_: data["data"]["children"][id_] for id_ in range(limit)}
data_dict = {}
for id_ in range(limit):
data_dict[id_] = {
item: data["data"]["children"][id_]["data"][item] for item in wanted_data
}
return data_dict
if __name__ == "__main__":
# If you get Error 429, that means you are rate limited.Try after some time
print(get_subreddit_data("learnpython", wanted_data=["title", "url", "selftext"]))