In today’s fast‑paced digital world, staying updated with the latest news, blog posts, and podcast episodes can feel overwhelming. An RSS (Really Simple Syndication) feed offers a lightweight, standardized way to pull content from multiple sources into a single, organized stream. Python, with its rich ecosystem of libraries, makes building a custom RSS feed parser and aggregator both straightforward and powerful. In this guide, we’ll explore why RSS still matters, walk through the most popular Python tools, and provide a step‑by‑step example that you can adapt for personal projects or production‑grade applications.
Why Build Your Own RSS Aggregator?
Before diving into code, it’s worth understanding the benefits of a custom aggregator:
- Control over content: Filter out unwanted topics, prioritize certain sources, or enrich items with additional metadata.
- Automation: Schedule regular fetches, store results in a database, and trigger notifications (email, Slack, or push alerts).
- Scalability: Combine dozens or hundreds of feeds into a single, searchable index without relying on third‑party services.
- Learning opportunity: Working with XML, HTTP requests, and asynchronous programming sharpens your Python skills.
Key Python Libraries for RSS Parsing
Several libraries simplify the process of fetching and parsing RSS/Atom feeds. Below is a quick comparison to help you choose the right tool for your needs.
feedparser
- Popularity: Over 10,000 stars on GitHub, widely used in the community.
- Features: Handles both RSS 2.0 and Atom, auto‑detects feed type, normalizes dates, and provides a simple dictionary‑like API.
- Pros: Minimal dependencies, easy to learn, great for quick prototypes.
- Cons: Synchronous only; may be slower when dealing with many feeds.
BeautifulSoup (with lxml)
- Popularity: A go‑to HTML/XML parser for many Python developers.
- Features: Full control over XML parsing, useful for non‑standard or malformed feeds.
- Pros: Flexible, works well with custom extraction logic.
- Cons: Requires more boilerplate than feedparser.
aiohttp + async‑feedparser
- Popularity: Growing in the async community.
- Features: Asynchronous HTTP requests combined with feed parsing, ideal for high‑throughput aggregators.
- Pros: Non‑blocking, scales to hundreds of feeds with minimal threads.
- Cons: Slightly steeper learning curve for newcomers to async programming.
Designing a Simple RSS Aggregator
Let’s build a minimal yet functional aggregator that:
- Downloads a list of RSS URLs.
- Parses each feed and extracts key fields (title, link, published date, summary).
- Filters items based on keywords.
- Stores the results in a SQLite database.
- Provides a command‑line interface to display the latest aggregated items.
Project Structure
rss_aggregator/
│
├─ aggregator.py # Core logic (fetch, parse, filter, store)
├─ models.py # SQLite ORM using SQLAlchemy or built‑in sqlite3
├─ config.py # Feed list and filter keywords
└─ requirements.txt # Dependencies
Step 1: Install Dependencies
Run the following command in your virtual environment:
pip install feedparser aiohttp aiosqlite
Step 2: Define the Feed List and Filters
Create config.py and populate it with a Python list. Keeping this data separate makes future updates painless.
# config.py
FEEDS = [
"https://realpython.com/atom.xml",
"https://news.ycombinator.com/rss",
"https://www.reddit.com/r/python/.rss",
]
# Simple keyword filter – only keep items containing any of these words.
KEYWORDS = ["python", "asyncio", "machine learning", "data science"]
Step 3: Database Model (SQLite)
For a lightweight solution, we’ll use aiosqlite to interact asynchronously with a SQLite file.
# models.py
import aiosqlite
DB_PATH = "feeds.db"
CREATE_TABLE_SQL = """
CREATE TABLE IF NOT EXISTS articles (
id INTEGER PRIMARY KEY AUTOINCREMENT,
feed_url TEXT,
title TEXT,
link TEXT,
published TEXT,
summary TEXT
);
"""
async def init_db():
async with aiosqlite.connect(DB_PATH) as db:
await db.execute(CREATE_TABLE_SQL)
await db.commit()
async def insert_article(db, article):
await db.execute(
"INSERT INTO articles (feed_url, title, link, published, summary) VALUES (?,?,?,?,?)",
(article["feed_url"], article["title"], article["link"],
article["published"], article["summary"])
)
Step 4: The Core Aggregator Logic
In aggregator.py, we’ll combine aiohttp for fetching and feedparser for parsing. The async approach keeps the I/O non‑blocking, which is essential when dealing with many feeds.
# aggregator.py
import asyncio
import aiohttp
import feedparser
import re
from datetime import datetime
from config import FEEDS, KEYWORDS
from models import init_db, insert_article, DB_PATH
def keyword_match(text):
"""Return True if any keyword appears in the given text (case‑insensitive)."""
pattern = re.compile("|".join(KEYWORDS), re.IGNORECASE)
return bool(pattern.search(text))
async def fetch_feed(session, url):
async with session.get(url, timeout=15) as response:
response.raise_for_status()
content = await response.text()
return feedparser.parse(content)
async def process_feed(session, db, url):
parsed = await fetch_feed(session, url)
for entry in parsed.entries:
# Normalise fields – different feeds use slightly different names.
title = entry.get("title", "No title")
link = entry.get("link", "")
published = entry.get("published", entry.get("updated", ""))
summary = entry.get("summary", entry.get("description", ""))
# Simple keyword filter
if not keyword_match(title + " " + summary):
continue
article = {
"feed_url": url,
"title": title,
"link": link,
"published": published,
"summary": summary,
}
await insert_article(db, article)
async def aggregate():
await init_db()
async with aiohttp.ClientSession() as session, aiosqlite.connect(DB_PATH) as db:
tasks = [process_feed(session, db, url) for url in FEEDS]
await asyncio.gather(*tasks)
await db.commit()
def pretty_date(date_str):
try:
dt = datetime.strptime(date_str[:25], "%a, %d %b %Y %H:%M:%S")
return dt.strftime("%Y-%m-%d %H:%M")
except Exception:
return date_str
async def show_latest(limit=10):
async with aiosqlite.connect(DB_PATH) as db:
async with db.execute(
"SELECT feed_url, title, link, published FROM articles ORDER BY id DESC LIMIT ?",
(limit,)
) as cursor:
rows = await cursor.fetchall()
for row in rows:
feed, title, link, pub = row
print(f"[{pretty_date(pub)}] {title}\\nSource: {feed}\\nLink: {link}\\n")
if __name__ == "__main__":
# Run the aggregator and then display the latest items.
asyncio.run(aggregate())
print("\\n=== Latest Aggregated Articles ===")
asyncio.run(show_latest())
Step 5: Running the Aggregator
Execute the script from the command line:
python aggregator.py
The first run creates feeds.db, fetches each RSS URL, filters items based on the keywords, and stores them. Subsequent runs will only add new entries, thanks to SQLite’s auto‑increment primary key.
Enhancing the Aggregator for Production Use
While the example above is functional, a real‑world service typically requires additional features:
- Duplicate detection: Store a hash of the article URL or GUID to avoid re‑inserting the same entry.
- Periodic scheduling: Use
cron,APScheduler, or a serverless function (AWS Lambda, Google Cloud Functions) to run the aggregator every hour. - Full‑text search: Integrate
SQLite FTS5or Elasticsearch to let users query the aggregated content. - Web front‑end: Build a Flask or FastAPI endpoint that returns JSON or renders a simple HTML page of the latest items.
- Authentication & rate limiting: Respect feed providers’ terms of service and avoid hammering servers with too‑frequent requests.
Sample Duplicate‑Check Logic
# Add a unique index on the link column
ALTER TABLE articles ADD UNIQUE (link);
# In insert_article, ignore duplicates
await db.execute(
"INSERT OR IGNORE INTO articles (feed_url, title, link, published, summary) VALUES (?,?,?,?,?)",
(...)
)
SEO Benefits of Publishing an RSS Aggregator
From an SEO perspective, a well‑structured aggregator page can attract organic traffic:
Leave a Reply