Python Rss Feed Parser And Aggregator

Written by

in

In today’s fast‑paced digital world, staying updated with the latest news, blog posts, and podcast episodes can feel overwhelming. An RSS (Really Simple Syndication) feed offers a lightweight, standardized way to pull content from multiple sources into a single, organized stream. Python, with its rich ecosystem of libraries, makes building a custom RSS feed parser and aggregator both straightforward and powerful. In this guide, we’ll explore why RSS still matters, walk through the most popular Python tools, and provide a step‑by‑step example that you can adapt for personal projects or production‑grade applications.

Why Build Your Own RSS Aggregator?

Before diving into code, it’s worth understanding the benefits of a custom aggregator:

  • Control over content: Filter out unwanted topics, prioritize certain sources, or enrich items with additional metadata.
  • Automation: Schedule regular fetches, store results in a database, and trigger notifications (email, Slack, or push alerts).
  • Scalability: Combine dozens or hundreds of feeds into a single, searchable index without relying on third‑party services.
  • Learning opportunity: Working with XML, HTTP requests, and asynchronous programming sharpens your Python skills.

Key Python Libraries for RSS Parsing

Several libraries simplify the process of fetching and parsing RSS/Atom feeds. Below is a quick comparison to help you choose the right tool for your needs.

feedparser

  • Popularity: Over 10,000 stars on GitHub, widely used in the community.
  • Features: Handles both RSS 2.0 and Atom, auto‑detects feed type, normalizes dates, and provides a simple dictionary‑like API.
  • Pros: Minimal dependencies, easy to learn, great for quick prototypes.
  • Cons: Synchronous only; may be slower when dealing with many feeds.

BeautifulSoup (with lxml)

  • Popularity: A go‑to HTML/XML parser for many Python developers.
  • Features: Full control over XML parsing, useful for non‑standard or malformed feeds.
  • Pros: Flexible, works well with custom extraction logic.
  • Cons: Requires more boilerplate than feedparser.

aiohttp + async‑feedparser

  • Popularity: Growing in the async community.
  • Features: Asynchronous HTTP requests combined with feed parsing, ideal for high‑throughput aggregators.
  • Pros: Non‑blocking, scales to hundreds of feeds with minimal threads.
  • Cons: Slightly steeper learning curve for newcomers to async programming.

Designing a Simple RSS Aggregator

Let’s build a minimal yet functional aggregator that:

  1. Downloads a list of RSS URLs.
  2. Parses each feed and extracts key fields (title, link, published date, summary).
  3. Filters items based on keywords.
  4. Stores the results in a SQLite database.
  5. Provides a command‑line interface to display the latest aggregated items.

Project Structure

rss_aggregator/
│
├─ aggregator.py          # Core logic (fetch, parse, filter, store)
├─ models.py              # SQLite ORM using SQLAlchemy or built‑in sqlite3
├─ config.py              # Feed list and filter keywords
└─ requirements.txt       # Dependencies

Step 1: Install Dependencies

Run the following command in your virtual environment:

pip install feedparser aiohttp aiosqlite

Step 2: Define the Feed List and Filters

Create config.py and populate it with a Python list. Keeping this data separate makes future updates painless.

# config.py
FEEDS = [
    "https://realpython.com/atom.xml",
    "https://news.ycombinator.com/rss",
    "https://www.reddit.com/r/python/.rss",
]

# Simple keyword filter – only keep items containing any of these words.
KEYWORDS = ["python", "asyncio", "machine learning", "data science"]

Step 3: Database Model (SQLite)

For a lightweight solution, we’ll use aiosqlite to interact asynchronously with a SQLite file.

# models.py
import aiosqlite

DB_PATH = "feeds.db"

CREATE_TABLE_SQL = """
CREATE TABLE IF NOT EXISTS articles (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    feed_url TEXT,
    title TEXT,
    link TEXT,
    published TEXT,
    summary TEXT
);
"""

async def init_db():
    async with aiosqlite.connect(DB_PATH) as db:
        await db.execute(CREATE_TABLE_SQL)
        await db.commit()

async def insert_article(db, article):
    await db.execute(
        "INSERT INTO articles (feed_url, title, link, published, summary) VALUES (?,?,?,?,?)",
        (article["feed_url"], article["title"], article["link"],
         article["published"], article["summary"])
    )

Step 4: The Core Aggregator Logic

In aggregator.py, we’ll combine aiohttp for fetching and feedparser for parsing. The async approach keeps the I/O non‑blocking, which is essential when dealing with many feeds.

# aggregator.py
import asyncio
import aiohttp
import feedparser
import re
from datetime import datetime

from config import FEEDS, KEYWORDS
from models import init_db, insert_article, DB_PATH

def keyword_match(text):
    """Return True if any keyword appears in the given text (case‑insensitive)."""
    pattern = re.compile("|".join(KEYWORDS), re.IGNORECASE)
    return bool(pattern.search(text))

async def fetch_feed(session, url):
    async with session.get(url, timeout=15) as response:
        response.raise_for_status()
        content = await response.text()
        return feedparser.parse(content)

async def process_feed(session, db, url):
    parsed = await fetch_feed(session, url)
    for entry in parsed.entries:
        # Normalise fields – different feeds use slightly different names.
        title = entry.get("title", "No title")
        link = entry.get("link", "")
        published = entry.get("published", entry.get("updated", ""))
        summary = entry.get("summary", entry.get("description", ""))

        # Simple keyword filter
        if not keyword_match(title + " " + summary):
            continue

        article = {
            "feed_url": url,
            "title": title,
            "link": link,
            "published": published,
            "summary": summary,
        }
        await insert_article(db, article)

async def aggregate():
    await init_db()
    async with aiohttp.ClientSession() as session, aiosqlite.connect(DB_PATH) as db:
        tasks = [process_feed(session, db, url) for url in FEEDS]
        await asyncio.gather(*tasks)
        await db.commit()

def pretty_date(date_str):
    try:
        dt = datetime.strptime(date_str[:25], "%a, %d %b %Y %H:%M:%S")
        return dt.strftime("%Y-%m-%d %H:%M")
    except Exception:
        return date_str

async def show_latest(limit=10):
    async with aiosqlite.connect(DB_PATH) as db:
        async with db.execute(
            "SELECT feed_url, title, link, published FROM articles ORDER BY id DESC LIMIT ?",
            (limit,)
        ) as cursor:
            rows = await cursor.fetchall()
            for row in rows:
                feed, title, link, pub = row
                print(f"[{pretty_date(pub)}] {title}\\nSource: {feed}\\nLink: {link}\\n")

if __name__ == "__main__":
    # Run the aggregator and then display the latest items.
    asyncio.run(aggregate())
    print("\\n=== Latest Aggregated Articles ===")
    asyncio.run(show_latest())

Step 5: Running the Aggregator

Execute the script from the command line:

python aggregator.py

The first run creates feeds.db, fetches each RSS URL, filters items based on the keywords, and stores them. Subsequent runs will only add new entries, thanks to SQLite’s auto‑increment primary key.

Enhancing the Aggregator for Production Use

While the example above is functional, a real‑world service typically requires additional features:

  • Duplicate detection: Store a hash of the article URL or GUID to avoid re‑inserting the same entry.
  • Periodic scheduling: Use cron, APScheduler, or a serverless function (AWS Lambda, Google Cloud Functions) to run the aggregator every hour.
  • Full‑text search: Integrate SQLite FTS5 or Elasticsearch to let users query the aggregated content.
  • Web front‑end: Build a Flask or FastAPI endpoint that returns JSON or renders a simple HTML page of the latest items.
  • Authentication & rate limiting: Respect feed providers’ terms of service and avoid hammering servers with too‑frequent requests.

Sample Duplicate‑Check Logic

# Add a unique index on the link column
ALTER TABLE articles ADD UNIQUE (link);

# In insert_article, ignore duplicates
await db.execute(
    "INSERT OR IGNORE INTO articles (feed_url, title, link, published, summary) VALUES (?,?,?,?,?)",
    (...)
)

SEO Benefits of Publishing an RSS Aggregator

From an SEO perspective, a well‑structured aggregator page can attract organic traffic:


Comments

Leave a Reply

Your email address will not be published. Required fields are marked *