Category: Uncategorized

  • Python Web Broken Link Checker Utility

    Keeping a website free of dead or misdirected URLs is essential for user experience, SEO rankings, and overall site credibility. A Python‑based broken link checker utility can automate the tedious task of crawling pages, validating hyperlinks, and reporting issues before they hurt your traffic. In this guide we’ll explore why link checking matters, outline the key features of a robust utility, walk through a step‑by‑step implementation, and share best‑practice tips for integrating the tool into your development workflow.

    Why a Broken Link Checker Is a Must‑Have for Modern Websites

    • Improves SEO: Search engines downgrade pages with many broken links, which can lower rankings.
    • Boosts user trust: Visitors quickly lose confidence when they encounter 404 errors.
    • Reduces bounce rate: A smooth navigation path keeps users engaged longer.
    • Supports accessibility: Screen readers announce broken links, creating a confusing experience for users with disabilities.
    • Automates maintenance: Manual checking is impractical for large sites; a script can run nightly or on every deployment.

    Core Features of an Effective Python Broken Link Checker

    • Recursive crawling of internal pages with a configurable depth limit.
    • Support for http, https, and protocol‑relative URLs.
    • Parallel request handling (threading or async) for speed.
    • Detection of common HTTP status codes (404, 500, 301/302 redirects, 429 throttling).
    • Whitelist/ignore patterns to skip external resources or known false positives.
    • Exportable reports in JSON, CSV, or plain‑text formats.
    • Easy integration with CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins).

    Building a Simple Yet Powerful Checker from Scratch

    1. Choose the Right Libraries

    For a balance of simplicity and performance, combine requests (or httpx for async) with beautifulsoup4 for HTML parsing. If you need high concurrency, aiohttp plus asyncio is a solid choice.

    2. Project Structure

    broken_link_checker/
    │
    ├─ checker.py          # Core crawling and validation logic
    ├─ utils.py            # Helper functions (URL normalization, logging)
    ├─ requirements.txt    # Dependencies
    └─ reports/
       └─ latest_report.json
    

    3. Sample Implementation (Synchronous Version)

    import re
    import json
    import queue
    import threading
    from urllib.parse import urljoin, urlparse
    
    import requests
    from bs4 import BeautifulSoup
    
    # ---------- Configuration ----------
    START_URL = "https://example.com"
    MAX_DEPTH = 3
    TIMEOUT = 10
    USER_AGENT = "PythonLinkChecker/1.0"
    IGNORE_PATTERNS = [r"^mailto:", r"^tel:", r"\.pdf$"]
    # ----------------------------------
    
    session = requests.Session()
    session.headers.update({"User-Agent": USER_AGENT})
    
    def should_ignore(link):
        return any(re.search(pat, link) for pat in IGNORE_PATTERNS)
    
    def is_internal(link):
        return urlparse(link).netloc in ("", urlparse(START_URL).netloc)
    
    def fetch(url):
        try:
            resp = session.get(url, timeout=TIMEOUT, allow_redirects=True)
            return resp.status_code, resp.url
        except requests.RequestException as e:
            return None, str(e)
    
    def extract_links(html, base):
        soup = BeautifulSoup(html, "html.parser")
        links = set()
        for tag in soup.find_all("a", href=True):
            href = tag["href"].strip()
            if should_ignore(href):
                continue
            absolute = urljoin(base, href)
            links.add(absolute.split("#")[0])  # Drop fragment
        return links
    
    def worker(q, results, visited):
        while True:
            try:
                url, depth = q.get_nowait()
            except queue.Empty:
                break
    
            if url in visited or depth > MAX_DEPTH:
                q.task_done()
                continue
    
            visited.add(url)
            status, final_url = fetch(url)
            results.append({"url": url, "status": status, "final_url": final_url})
    
            if status and 200 <= status < 400 and is_internal(final_url):
                try:
                    resp = session.get(final_url, timeout=TIMEOUT)
                    child_links = extract_links(resp.text, final_url)
                    for link in child_links:
                        q.put((link, depth + 1))
                except requests.RequestException:
                    pass
    
            q.task_done()
    
    def run_checker():
        q = queue.Queue()
        q.put((START_URL, 0))
        results = []
        visited = set()
        threads = []
    
        for _ in range(8):  # Number of worker threads
            t = threading.Thread(target=worker, args=(q, results, visited))
            t.start()
            threads.append(t)
    
        q.join()
        for t in threads:
            t.join()
    
        return results
    
    if __name__ == "__main__":
        report = run_checker()
        with open("reports/latest_report.json", "w", encoding="utf-8") as f:
            json.dump(report, f, indent=2)
        print(f"✅ Scan complete – {len(report)} URLs processed.")
    

    4. Key Points in the Code

    • Queue‑based crawling: Guarantees breadth‑first traversal and easy depth control.
    • Thread pool: Eight workers provide a good speed‑up without overwhelming the target server.
    • Ignore patterns: Regular expressions prevent false positives on mailto links, PDFs, or JavaScript URLs.
    • Result recording: Each entry stores the original URL, HTTP status, and final redirected URL.

    Scaling Up with AsyncIO for Massive Sites

    If you need to scan thousands of pages, switch to aiohttp and asyncio.Semaphore to limit concurrent connections. Below is a concise async skeleton you can drop into checker.py:

    import asyncio
    import aiohttp
    from bs4 import BeautifulSoup
    from urllib.parse import urljoin, urlparse
    
    SEM = asyncio.Semaphore(20)  # Max 20 concurrent requests
    
    async def fetch(session, url):
        async with SEM:
            try:
                async with session.get(url, timeout=TIMEOUT) as resp:
                    text = await resp.text()
                    return resp.status, str(resp.url), text
            except Exception as e:
                return None, str(e), ""
    
    async def crawl(url, depth, session, visited, results):
        if url in visited or depth > MAX_DEPTH:
            return
        visited.add(url)
    
        status, final_url, html = await fetch(session, url)
        results.append({"url": url, "status": status, "final_url": final_url})
    
        if status and 200 <= status < 400 and is_internal(final_url):
            soup = BeautifulSoup(html, "html.parser")
            for tag in soup.find_all("a", href=True):
                link = urljoin(final_url, tag["href"]).split("#")[0]
                if not should_ignore(link):
                    await crawl(link, depth + 1, session, visited, results)
    
    async def main():
        async with aiohttp.ClientSession(headers={"User-Agent": USER_AGENT}) as session:
            visited = set()
            results = []
            await crawl(START_URL, 0, session, visited, results)
            # Save JSON report...
            ...
    
    if __name__ == "__main__":
        asyncio.run(main())
    

    Integrating the Checker Into Your CI/CD Pipeline

    Automated link validation should be part of every deployment. Here’s a quick checklist for a GitHub Actions workflow:

    1. Create a requirements.txt with requests, beautifulsoup4, and optionally aiohttp.
    2. Add a job that installs dependencies and runs python checker.py.
    3. Fail the build if the JSON report contains any entry with status >= 400.
    4. Upload the report as an artifact for later review.
    name: Link Check
    on: [push, pull_request]
    
    jobs:
      broken-links:
        runs-on: ubuntu-latest
        steps:
          - uses: actions/checkout@v3
          - name: Set up Python
            uses: actions/setup-python@v4
            with:
              python-version: "3.12"
          - name: Install deps
            run: pip install -r requirements.txt
          - name: Run checker
            run: python checker.py
          - name: Fail on broken links
            run: |
              python -c "
    import json, sys
    with open('reports/latest_report.json') as f:
        data = json.load(f)
    bad = [r for r in data if r['status'] and r['status'] >= 400]
    if bad:
        print('Broken links found:', len(bad))
        sys.exit(1)
    "
          - name: Upload report
            uses: actions/upload-artifact@v3
            with:
              name: link-report
              path: reports/latest_report.json
    

    Advanced Tips for a Production‑Ready Utility

    Handle Rate Limiting Gracefully

    • Detect 429 Too Many Requests and respect the Retry-After header.
    • Implement exponential back‑off for repeated failures.

    Cache Results to Speed Up Re‑scans

    Store the last known status of each URL in a lightweight SQLite database. On subsequent runs, skip URLs that haven’t changed

  • Python Rss Feed Parser And Aggregator

    In today’s fast‑paced digital world, staying updated with the latest news, blog posts, and podcast episodes can feel overwhelming. An RSS (Really Simple Syndication) feed offers a lightweight, standardized way to pull content from multiple sources into a single, organized stream. Python, with its rich ecosystem of libraries, makes building a custom RSS feed parser and aggregator both straightforward and powerful. In this guide, we’ll explore why RSS still matters, walk through the most popular Python tools, and provide a step‑by‑step example that you can adapt for personal projects or production‑grade applications.

    Why Build Your Own RSS Aggregator?

    Before diving into code, it’s worth understanding the benefits of a custom aggregator:

    • Control over content: Filter out unwanted topics, prioritize certain sources, or enrich items with additional metadata.
    • Automation: Schedule regular fetches, store results in a database, and trigger notifications (email, Slack, or push alerts).
    • Scalability: Combine dozens or hundreds of feeds into a single, searchable index without relying on third‑party services.
    • Learning opportunity: Working with XML, HTTP requests, and asynchronous programming sharpens your Python skills.

    Key Python Libraries for RSS Parsing

    Several libraries simplify the process of fetching and parsing RSS/Atom feeds. Below is a quick comparison to help you choose the right tool for your needs.

    feedparser

    • Popularity: Over 10,000 stars on GitHub, widely used in the community.
    • Features: Handles both RSS 2.0 and Atom, auto‑detects feed type, normalizes dates, and provides a simple dictionary‑like API.
    • Pros: Minimal dependencies, easy to learn, great for quick prototypes.
    • Cons: Synchronous only; may be slower when dealing with many feeds.

    BeautifulSoup (with lxml)

    • Popularity: A go‑to HTML/XML parser for many Python developers.
    • Features: Full control over XML parsing, useful for non‑standard or malformed feeds.
    • Pros: Flexible, works well with custom extraction logic.
    • Cons: Requires more boilerplate than feedparser.

    aiohttp + async‑feedparser

    • Popularity: Growing in the async community.
    • Features: Asynchronous HTTP requests combined with feed parsing, ideal for high‑throughput aggregators.
    • Pros: Non‑blocking, scales to hundreds of feeds with minimal threads.
    • Cons: Slightly steeper learning curve for newcomers to async programming.

    Designing a Simple RSS Aggregator

    Let’s build a minimal yet functional aggregator that:

    1. Downloads a list of RSS URLs.
    2. Parses each feed and extracts key fields (title, link, published date, summary).
    3. Filters items based on keywords.
    4. Stores the results in a SQLite database.
    5. Provides a command‑line interface to display the latest aggregated items.

    Project Structure

    rss_aggregator/
    │
    ├─ aggregator.py          # Core logic (fetch, parse, filter, store)
    ├─ models.py              # SQLite ORM using SQLAlchemy or built‑in sqlite3
    ├─ config.py              # Feed list and filter keywords
    └─ requirements.txt       # Dependencies
    

    Step 1: Install Dependencies

    Run the following command in your virtual environment:

    pip install feedparser aiohttp aiosqlite

    Step 2: Define the Feed List and Filters

    Create config.py and populate it with a Python list. Keeping this data separate makes future updates painless.

    # config.py
    FEEDS = [
        "https://realpython.com/atom.xml",
        "https://news.ycombinator.com/rss",
        "https://www.reddit.com/r/python/.rss",
    ]
    
    # Simple keyword filter – only keep items containing any of these words.
    KEYWORDS = ["python", "asyncio", "machine learning", "data science"]
    

    Step 3: Database Model (SQLite)

    For a lightweight solution, we’ll use aiosqlite to interact asynchronously with a SQLite file.

    # models.py
    import aiosqlite
    
    DB_PATH = "feeds.db"
    
    CREATE_TABLE_SQL = """
    CREATE TABLE IF NOT EXISTS articles (
        id INTEGER PRIMARY KEY AUTOINCREMENT,
        feed_url TEXT,
        title TEXT,
        link TEXT,
        published TEXT,
        summary TEXT
    );
    """
    
    async def init_db():
        async with aiosqlite.connect(DB_PATH) as db:
            await db.execute(CREATE_TABLE_SQL)
            await db.commit()
    
    async def insert_article(db, article):
        await db.execute(
            "INSERT INTO articles (feed_url, title, link, published, summary) VALUES (?,?,?,?,?)",
            (article["feed_url"], article["title"], article["link"],
             article["published"], article["summary"])
        )
    

    Step 4: The Core Aggregator Logic

    In aggregator.py, we’ll combine aiohttp for fetching and feedparser for parsing. The async approach keeps the I/O non‑blocking, which is essential when dealing with many feeds.

    # aggregator.py
    import asyncio
    import aiohttp
    import feedparser
    import re
    from datetime import datetime
    
    from config import FEEDS, KEYWORDS
    from models import init_db, insert_article, DB_PATH
    
    def keyword_match(text):
        """Return True if any keyword appears in the given text (case‑insensitive)."""
        pattern = re.compile("|".join(KEYWORDS), re.IGNORECASE)
        return bool(pattern.search(text))
    
    async def fetch_feed(session, url):
        async with session.get(url, timeout=15) as response:
            response.raise_for_status()
            content = await response.text()
            return feedparser.parse(content)
    
    async def process_feed(session, db, url):
        parsed = await fetch_feed(session, url)
        for entry in parsed.entries:
            # Normalise fields – different feeds use slightly different names.
            title = entry.get("title", "No title")
            link = entry.get("link", "")
            published = entry.get("published", entry.get("updated", ""))
            summary = entry.get("summary", entry.get("description", ""))
    
            # Simple keyword filter
            if not keyword_match(title + " " + summary):
                continue
    
            article = {
                "feed_url": url,
                "title": title,
                "link": link,
                "published": published,
                "summary": summary,
            }
            await insert_article(db, article)
    
    async def aggregate():
        await init_db()
        async with aiohttp.ClientSession() as session, aiosqlite.connect(DB_PATH) as db:
            tasks = [process_feed(session, db, url) for url in FEEDS]
            await asyncio.gather(*tasks)
            await db.commit()
    
    def pretty_date(date_str):
        try:
            dt = datetime.strptime(date_str[:25], "%a, %d %b %Y %H:%M:%S")
            return dt.strftime("%Y-%m-%d %H:%M")
        except Exception:
            return date_str
    
    async def show_latest(limit=10):
        async with aiosqlite.connect(DB_PATH) as db:
            async with db.execute(
                "SELECT feed_url, title, link, published FROM articles ORDER BY id DESC LIMIT ?",
                (limit,)
            ) as cursor:
                rows = await cursor.fetchall()
                for row in rows:
                    feed, title, link, pub = row
                    print(f"[{pretty_date(pub)}] {title}\\nSource: {feed}\\nLink: {link}\\n")
    
    if __name__ == "__main__":
        # Run the aggregator and then display the latest items.
        asyncio.run(aggregate())
        print("\\n=== Latest Aggregated Articles ===")
        asyncio.run(show_latest())
    

    Step 5: Running the Aggregator

    Execute the script from the command line:

    python aggregator.py

    The first run creates feeds.db, fetches each RSS URL, filters items based on the keywords, and stores them. Subsequent runs will only add new entries, thanks to SQLite’s auto‑increment primary key.

    Enhancing the Aggregator for Production Use

    While the example above is functional, a real‑world service typically requires additional features:

    • Duplicate detection: Store a hash of the article URL or GUID to avoid re‑inserting the same entry.
    • Periodic scheduling: Use cron, APScheduler, or a serverless function (AWS Lambda, Google Cloud Functions) to run the aggregator every hour.
    • Full‑text search: Integrate SQLite FTS5 or Elasticsearch to let users query the aggregated content.
    • Web front‑end: Build a Flask or FastAPI endpoint that returns JSON or renders a simple HTML page of the latest items.
    • Authentication & rate limiting: Respect feed providers’ terms of service and avoid hammering servers with too‑frequent requests.

    Sample Duplicate‑Check Logic

    # Add a unique index on the link column
    ALTER TABLE articles ADD UNIQUE (link);
    
    # In insert_article, ignore duplicates
    await db.execute(
        "INSERT OR IGNORE INTO articles (feed_url, title, link, published, summary) VALUES (?,?,?,?,?)",
        (...)
    )
    

    SEO Benefits of Publishing an RSS Aggregator

    From an SEO perspective, a well‑structured aggregator page can attract organic traffic:


  • Python News Aggregator Web App Project

    In today’s fast‑paced digital world, staying on top of the latest headlines can feel like chasing a moving target. A Python‑powered news aggregator web app solves that problem by pulling stories from multiple sources, normalizing the data, and presenting it in a clean, customizable interface. Whether you’re a beginner looking for a hands‑on project or a seasoned developer aiming to showcase your full‑stack skills, building a news aggregator with Python offers a perfect blend of API integration, data processing, and web development. In this guide we’ll walk through everything you need to know to plan, code, and launch a robust news aggregator that not only works flawlessly but also ranks well in search engines.

    What Is a News Aggregator and Why Build One?

    A news aggregator is a platform that collects articles, blog posts, and other content from various sources—RSS feeds, public APIs, or web‑scraped pages—and displays them in a single, unified feed. The main benefits of creating your own aggregator include:

    • Personalization: Tailor the content to specific topics, languages, or regions.
    • Learning Opportunity: Gain hands‑on experience with HTTP requests, JSON handling, and database design.
    • Portfolio Boost: Demonstrate full‑stack capabilities to potential employers or clients.
    • Monetization Potential: Add ads, premium features, or affiliate links for revenue.

    Why Choose Python for This Project?

    Python’s readability, extensive libraries, and vibrant community make it an ideal choice for a news aggregator. Here are the top reasons to pick Python:

    • Rich Ecosystem: Libraries like requests, feedparser, and BeautifulSoup simplify API calls and web scraping.
    • Fast Prototyping: Write less boilerplate code compared to compiled languages.
    • Scalable Frameworks: Flask and Django provide lightweight and heavyweight options for the backend.
    • SEO‑Friendly: Python’s templating engines (Jinja2, Django templates) let you generate clean HTML that search bots love.

    Core Features Every Good Aggregator Should Have

    Essential Functionality

    • Source management (add, edit, remove RSS feeds or API endpoints).
    • Automatic fetching and scheduling (e.g., every 30 minutes).
    • Article parsing (title, description, author, publish date, image).
    • Keyword‑based filtering and categorization.
    • User‑friendly UI with pagination or infinite scroll.

    Nice‑to‑Have Enhancements

    • User authentication and saved preferences.
    • Full‑text search powered by Whoosh or Elasticsearch.
    • Responsive design using Bootstrap or Tailwind CSS.
    • Social sharing buttons and comment integration.
    • Analytics dashboard to track most‑read topics.

    Tech Stack Overview

    Language Python 3.11+
    Web Framework Flask (lightweight) or Django (full‑featured)
    Database SQLite for prototyping, PostgreSQL for production
    Task Scheduler Celery with Redis broker or APScheduler
    Front‑End HTML5, CSS3, JavaScript (vanilla or Vue/React)
    Hosting Heroku, Render, or AWS Elastic Beanstalk

    Step‑by‑Step Development Guide

    1. Set Up the Project Structure

    news_aggregator/
    │
    ├─ app/
    │   ├─ __init__.py
    │   ├─ routes.py
    │   ├─ models.py
    │   └─ utils.py
    │
    ├─ templates/
    │   └─ index.html
    │
    ├─ static/
    │   ├─ css/
    │   └─ js/
    │
    ├─ requirements.txt
    └─ run.py
    

    Use a virtual environment and install the core dependencies:

    python -m venv venv
    source venv/bin/activate  # Windows: venv\Scripts\activate
    pip install flask sqlalchemy requests feedparser APScheduler
    

    2. Define the Database Model

    from datetime import datetime
    from flask_sqlalchemy import SQLAlchemy
    
    db = SQLAlchemy()
    
    class Article(db.Model):
        id = db.Column(db.Integer, primary_key=True)
        source = db.Column(db.String(120), nullable=False)
        title = db.Column(db.String(255), nullable=False)
        url = db.Column(db.String(500), nullable=False, unique=True)
        summary = db.Column(db.Text)
        published = db.Column(db.DateTime, default=datetime.utcnow)
        image_url = db.Column(db.String(500))
    

    This simple schema stores the most important metadata for each article while keeping the table lightweight for fast queries.

    3. Fetch News from External Sources

    Below is a reusable function that pulls data from an RSS feed using feedparser and stores new articles in the database.

    import feedparser
    from .models import Article, db
    
    def fetch_rss(feed_url):
        feed = feedparser.parse(feed_url)
        for entry in feed.entries:
            # Avoid duplicates
            if Article.query.filter_by(url=entry.link).first():
                continue
    
            article = Article(
                source=feed.feed.title,
                title=entry.title,
                url=entry.link,
                summary=getattr(entry, 'summary', ''),
                published=getattr(entry, 'published_parsed', None),
                image_url=entry.media_content[0]['url'] if 'media_content' in entry else ''
            )
            db.session.add(article)
        db.session.commit()
    

    For JSON‑based APIs (e.g., NewsAPI.org), replace feedparser with requests.get() and parse the JSON response accordingly.

    4. Schedule Automatic Updates

    Using APScheduler, you can run fetch_rss every 30 minutes without leaving the server process.

    from apscheduler.schedulers.background import BackgroundScheduler
    from .utils import fetch_rss
    
    def start_scheduler(app):
        scheduler = BackgroundScheduler()
        scheduler.add_job(
            func=lambda: fetch_rss('https://rss.cnn.com/rss/edition.rss'),
            trigger='interval',
            minutes=30,
            id='cnn_feed'
        )
        scheduler.start()
        # Shut down gracefully when the app exits
        app.teardown_appcontext(lambda exception: scheduler.shutdown())
    

    5. Build the Flask Routes

    from flask import Flask, render_template, request
    from .models import db, Article
    from .utils import start_scheduler
    
    def create_app():
        app = Flask(__name__)
        app.config['SQLALCHEMY_DATABASE_URI'] = 'sqlite:///news.db'
        db.init_app(app)
    
        @app.route('/')
        def index():
            page = request.args.get('page', 1, type=int)
            articles = Article.query.order_by(Article.published.desc()).paginate(page, 10, False)
            return render_template('index.html', articles=articles)
    
        with app.app_context():
            db.create_all()
            start_scheduler(app)
    
        return app
    

    6. Design a Clean Front‑End

    Save the following as templates/index.html. It uses Bootstrap CDN for quick styling and includes SEO‑friendly meta tags.

    <!DOCTYPE html>
    <html lang="en">
    <head>
    <meta charset="UTF-8">
    <meta name="viewport" content="width=device-width, initial-scale=1.0">
    <meta name="description" content="Live Python‑powered news aggregator pulling top stories from RSS feeds and APIs.">
    <title>Python News Aggregator</title>
    <link href="https://cdn.jsdelivr.net/npm/bootstrap@5.3.3/dist/css/bootstrap.min.css" rel="stylesheet">
    </head>
    <body class="bg-light">
    <div class="container py-4">
    <h2 class="mb-4">Latest Headlines</h2>
    <div class="row">
    {% for article in articles.items %}
    <div class="col-md-6 mb-3">
    <div class="card h-100">
    {% if article.image_url %}
    <img src="{{ article.image_url }}" class="card-img-top" alt="Article image">
    {% endif %}
    <div class="card-body">
    <h5 class="card-title">
    <a href="{{ article.url }}" target="_blank" rel="noopener">{{ article.title }}</a>
    </h5>
    <p class="card-text">{{ article.summary|truncate(150) }}</p>
    </div>
    &

  • Python E-Commerce Price Drop Alert Tool

    In the fast‑moving world of online shopping, a few dollars can make the difference between a purchase and a missed opportunity. That’s why savvy shoppers rely on price‑drop alerts to snag the best deals the moment they happen. If you’re a developer or a hobbyist looking to automate this process, a Python e‑commerce price drop alert tool is the perfect solution. In this guide we’ll explore why price‑drop alerts matter, outline the essential features of a robust tool, and walk you through a step‑by‑step implementation that you can customize for any online store.

    Why Build a Python Price‑Drop Alert Tool?

    Before diving into the code, let’s understand the value this tool brings:

    • Instant Savings: Get notified the moment a product’s price falls below your target, ensuring you never overpay.
    • Time Efficiency: Automate the tedious task of manually checking product pages multiple times a day.
    • Competitive Edge: Stay ahead of other shoppers, especially during flash sales and limited‑time promotions.
    • Scalability: Monitor dozens or even hundreds of items simultaneously with minimal overhead.

    Core Features of an Effective Alert System

    1. Reliable Web Scraping

    Extract product titles, current prices, and stock status from e‑commerce sites using libraries such as requests, BeautifulSoup, or Selenium for JavaScript‑heavy pages.

    2. Intelligent Price Tracking

    Store historical price data to detect drops, calculate percentage changes, and avoid false positives caused by temporary price fluctuations.

    3. Customizable Notification Channels

    Send alerts via email, SMS, push notifications, or messaging platforms (Telegram, Discord). Using smtplib for email or APIs like Twilio for SMS keeps the system flexible.

    4. Persistent Storage

    Persist product URLs, target prices, and last‑known prices in a lightweight database (SQLite) or a CSV file for quick prototyping.

    5. Scheduling & Automation

    Run the scraper at regular intervals using cron, APScheduler, or cloud‑based schedulers (AWS Lambda, Google Cloud Functions).

    Step‑by‑Step Guide: Building Your Own Tool

    Step 1: Set Up the Project Environment

    mkdir price-drop-alert
    cd price-drop-alert
    python -m venv venv
    source venv/bin/activate   # On Windows: venv\Scripts\activate
    pip install requests beautifulsoup4 selenium apscheduler pandas
    

    Step 2: Choose the Right Scraping Method

    For static pages, requests + BeautifulSoup is fast and lightweight. For dynamic content rendered by JavaScript, use Selenium with a headless browser.

    Static Scraper Example

    import requests
    from bs4 import BeautifulSoup
    
    def fetch_price(url):
        headers = {
            "User-Agent": "Mozilla/5.0 (compatible; PriceDropBot/1.0)"
        }
        response = requests.get(url, headers=headers, timeout=10)
        soup = BeautifulSoup(response.text, "html.parser")
        # Example: Amazon price span with id="priceblock_ourprice"
        price_tag = soup.select_one("#priceblock_ourprice")
        if price_tag:
            price_text = price_tag.get_text().strip().replace("$", "")
            return float(price_text)
        return None
    

    Dynamic Scraper Skeleton

    from selenium import webdriver
    from selenium.webdriver.chrome.options import Options
    
    def fetch_price_dynamic(url):
        options = Options()
        options.add_argument("--headless")
        driver = webdriver.Chrome(options=options)
        driver.get(url)
        # Adjust selector based on the site
        price_elem = driver.find_element_by_css_selector(".price")
        price = float(price_elem.text.replace("$", ""))
        driver.quit()
        return price
    

    Step 3: Store and Compare Prices

    We’ll use pandas to manage a CSV file called products.csv with columns: url, target_price, last_price.

    import pandas as pd
    
    CSV_PATH = "products.csv"
    
    def load_products():
        return pd.read_csv(CSV_PATH)
    
    def save_products(df):
        df.to_csv(CSV_PATH, index=False)
    
    def check_price_drop(product):
        current_price = fetch_price(product["url"])
        if current_price is None:
            return None
        if current_price <= product["target_price"] and current_price < product["last_price"]:
            return current_price
        return None
    

    Step 4: Send Notifications

    Below is a simple email alert using smtplib. Replace placeholders with your SMTP credentials.

    import smtplib
    from email.mime.text import MIMEText
    
    SMTP_SERVER = "smtp.gmail.com"
    SMTP_PORT = 587
    SMTP_USER = "your.email@gmail.com"
    SMTP_PASS = "your_app_password"
    
    def send_email(subject, body, to=SMTP_USER):
        msg = MIMEText(body, "plain")
        msg["Subject"] = subject
        msg["From"] = SMTP_USER
        msg["To"] = to
    
        with smtplib.SMTP(SMTP_SERVER, SMTP_PORT) as server:
            server.starttls()
            server.login(SMTP_USER, SMTP_PASS)
            server.send_message(msg)
    

    Step 5: Orchestrate the Workflow with APScheduler

    from apscheduler.schedulers.blocking import BlockingScheduler
    
    def monitor():
        df = load_products()
        for idx, row in df.iterrows():
            new_price = check_price_drop(row)
            if new_price is not None:
                subject = f"💰 Price Drop Alert: {row['url']}"
                body = (f"The price has fallen to ${new_price:.2f}!\n"
                        f"Target price: ${row['target_price']:.2f}\n"
                        f"Buy now: {row['url']}")
                send_email(subject, body)
                df.at[idx, "last_price"] = new_price
        save_products(df)
    
    scheduler = BlockingScheduler()
    scheduler.add_job(monitor, "interval", minutes=30)  # Adjust frequency as needed
    scheduler.start()
    

    Handling Real‑World Challenges

    Anti‑Scraping Measures

    • Rotate User‑Agents: Randomly select a realistic User‑Agent string for each request.
    • Use Proxies: Services like ScraperAPI or residential proxy pools help avoid IP bans.
    • Respect robots.txt: While not legally binding, honoring robots.txt reduces the risk of being blocked.

    Dealing with Price Variability

    Some sites show different prices based on location or logged‑in status. To improve accuracy:

    1. Set a consistent Accept-Language header.
    2. Maintain a persistent session with cookies.
    3. If needed, log in programmatically using Selenium before scraping.

    Scaling Up

    When monitoring hundreds of products, consider these upgrades:

    • Database: Switch from CSV to PostgreSQL or MySQL for faster queries.
    • Task Queue: Use Celery with Redis/RabbitMQ to parallelize scraping jobs.
    • Containerization: Deploy the entire stack in Docker for consistent environments.

    Deploying the Tool to the Cloud

    Running the script on a local machine works for testing, but a cloud deployment ensures 24/7 monitoring.

    Option 1: Heroku (Free Tier)

    1. Create a Procfile with worker: python monitor.py.
    2. Push the repository to Heroku and enable the worker dyno.
    3. Use Heroku Scheduler to trigger the script every 30 minutes.

    Option 2: AWS Lambda + CloudWatch

    1. Package the code and dependencies into a ZIP file.
    2. Create a Lambda function with a Python 3.11 runtime.
    3. Set a CloudWatch Events rule to invoke the function on a schedule.

    Option 3: Docker + VPS

    # Dockerfile
    FROM python:3.11-slim
    WORKDIR /app
    COPY requirements.txt .
    RUN pip install -r requirements.txt
    COPY . .
    CMD ["python", "monitor.py"]
    

    Build and run the container on any VPS, then use cron inside the container or host‑level scheduling.

    SEO Tips for Your Blog Post

    • Keyword Placement: Include “Python e‑commerce price drop alert tool” in the first 100 words, in at least one <h2>, and naturally throughout the content.
    • Meta Description: Write a concise 150‑character summary containing the main keyword.
    • <
  • Python Real-Time Stock Price Scraper

    In the fast‑paced world of finance, having access to up‑to‑the‑second stock prices can be the difference between a profitable trade and a missed opportunity. Python, with its rich ecosystem of libraries, makes it surprisingly easy to build a real‑time stock price scraper that pulls live data from websites, APIs, or even WebSocket streams. In this guide, you’ll learn why real‑time data matters, how to choose the right source, which Python tools to use, and a step‑by‑step walkthrough to create a robust, production‑ready scraper—all while keeping SEO best practices in mind.

    Why Real‑Time Stock Data Matters

    Investors, algorithmic traders, and financial analysts rely on live market information to make split‑second decisions. Unlike end‑of‑day summaries, real‑time data provides:

    • Immediate price movements that reveal market sentiment.
    • The ability to execute automated strategies such as arbitrage or momentum trading.
    • Up‑to‑date risk management metrics for portfolio monitoring.
    • Insights for news‑driven trading, where a headline can shift prices within seconds.

    Benefits for Traders and Developers

    When you build a Python real‑time stock price scraper, you gain:

    • Customization – Tailor the data fields (price, volume, bid/ask) to your exact needs.
    • Cost efficiency – Avoid expensive data‑feed subscriptions by leveraging free web sources.
    • Scalability – Python’s async capabilities let you monitor dozens or hundreds of symbols simultaneously.
    • Learning opportunity – Master web‑scraping, API integration, and data‑pipeline design in a single project.

    Choosing the Right Data Source

    Before writing a single line of code, decide where you’ll pull the stock prices from. The two main options are free APIs and direct web scraping.

    Free APIs vs. Web Scraping

    • APIs (e.g., Alpha Vantage, IEX Cloud, Finnhub) usually provide structured JSON, clear rate limits, and documentation, making them developer‑friendly.
    • Web scraping extracts data directly from a website’s HTML or JavaScript. It can be more flexible (access to sites without public APIs) but requires handling dynamic content and respecting legal constraints.

    Legal and Ethical Considerations

    Always read a site’s robots.txt and terms of service. Scraping financial data for commercial use without permission may violate copyright or licensing agreements. When in doubt, opt for a reputable API or seek explicit permission.

    Core Tools for a Python Stock Scraper

    The Python ecosystem offers several libraries that make real‑time scraping straightforward:

    • requests – Simple HTTP client for REST APIs and static pages.
    • BeautifulSoup – Parses HTML and XML, perfect for extracting table data.
    • Selenium – Automates a web browser to handle JavaScript‑rendered pages.
    • websocket‑client – Connects to live WebSocket streams when available.
    • pandas – Organizes, cleans, and analyses time‑series data.
    • APScheduler or cron – Schedules periodic scraping tasks.

    Step‑by‑Step Guide: Building a Real‑Time Scraper

    1. Setting Up the Environment

    Start with a virtual environment to keep dependencies isolated:

    python -m venv stock-scraper-env
    source stock-scraper-env/bin/activate  # Linux/macOS
    .\stock-scraper-env\Scripts\activate   # Windows
    
    pip install requests beautifulsoup4 selenium pandas APScheduler websocket-client
    

    2. Fetching Data with requests

    If the target site offers a simple JSON endpoint, a quick GET request may be enough:

    import requests
    
    def fetch_price(symbol):
        url = f"https://api.example.com/quote/{symbol}"
        response = requests.get(url, timeout=5)
        response.raise_for_status()
        data = response.json()
        return data["price"]
    

    3. Parsing HTML with BeautifulSoup

    For sites that render prices in static HTML tables, use BeautifulSoup to locate the element:

    from bs4 import BeautifulSoup
    
    def scrape_price(symbol):
        url = f"https://www.example.com/stocks/{symbol}"
        html = requests.get(url).text
        soup = BeautifulSoup(html, "html.parser")
        price_tag = soup.find("div", class_="price-value")
        return float(price_tag.text.replace("$", ""))
    

    4. Handling JavaScript‑Rendered Pages with Selenium

    When the price appears only after JavaScript execution, Selenium drives a headless browser:

    from selenium import webdriver
    from selenium.webdriver.chrome.options import Options
    
    def selenium_price(symbol):
        options = Options()
        options.add_argument("--headless")
        driver = webdriver.Chrome(options=options)
        driver.get(f"https://www.example.com/quote/{symbol}")
    
        # Wait for the price element to load (simple sleep for demo)
        driver.implicitly_wait(5)
        price_elem = driver.find_element("css selector", ".price-value")
        price = float(price_elem.text.replace("$", ""))
        driver.quit()
        return price
    

    5. Streaming Updates via WebSocket (Optional)

    Some platforms expose a live feed through WebSocket. Here’s a minimal example using websocket-client:

    import json
    from websocket import create_connection
    
    def stream_prices(symbol):
        ws = create_connection("wss://stream.example.com/realtime")
        ws.send(json.dumps({"type": "subscribe", "symbol": symbol}))
        while True:
            result = ws.recv()
            data = json.loads(result)
            print(f"{symbol}: {data['price']} at {data['timestamp']}")
    

    6. Storing and Analyzing Data with pandas

    Collect the scraped prices into a DataFrame for quick analysis:

    import pandas as pd
    from datetime import datetime
    
    prices = []
    
    def record_price(symbol):
        price = fetch_price(symbol)  # or any of the methods above
        prices.append({"symbol": symbol, "price": price, "time": datetime.utcnow()})
    
    def to_dataframe():
        df = pd.DataFrame(prices)
        df.set_index("time", inplace=True)
        return df
    

    7. Scheduling with APScheduler or Cron

    To keep the scraper running continuously, schedule it to execute every few seconds:

    from apscheduler.schedulers.blocking import BlockingScheduler
    
    scheduler = BlockingScheduler()
    scheduler.add_job(lambda: record_price("AAPL"), "interval", seconds=5)
    scheduler.add_job(lambda: record_price("GOOGL"), "interval", seconds=5)
    
    try:
        scheduler.start()
    except (KeyboardInterrupt, SystemExit):
        pass
    

    Common Pitfalls and How to Avoid Them

    Rate Limiting and IP Blocking

    • Respect the source’s rate limits. Use time.sleep() or a scheduler to throttle requests.
    • Rotate proxy IPs or employ services like scraperapi.com for high‑volume scraping.
    • Include realistic User‑Agent headers to mimic a regular browser.

    Data Accuracy and Stale Prices

    Always verify that the timestamp of the retrieved price matches the current market session. Compare against a known reliable API periodically to detect drift.

    Handling Unexpected HTML Changes

  • Python Automated Job Application Scraper

    Imagine waking up every morning to a fresh list of tailored job openings, each ready to be applied to with a single click. Thanks to Python’s powerful web‑scraping libraries and automation tools, this vision can become a reality. In this guide we’ll walk you through everything you need to build a reliable, SEO‑friendly Python automated job application scraper—from choosing the right libraries to handling anti‑scraping defenses and scaling your solution for daily use.

    Why Automate Job Applications with Python?

    Job hunting is time‑consuming. A typical job seeker spends hours each week browsing multiple job boards, copying job descriptions, and manually filling out application forms. Automation offers several compelling benefits:

    • Speed: Scrape hundreds of listings in minutes and submit applications instantly.
    • Precision: Filter jobs by location, salary, keywords, and experience level without human error.
    • Consistency: Ensure every application follows the same format, increasing your professional image.
    • Scalability: Expand your search to niche boards and company career pages without extra effort.

    When optimized for search engines, a well‑structured blog post on this topic also attracts developers looking for practical examples, boosting organic traffic and establishing authority in the Python automation niche.

    Key Components of a Job Scraper

    1. HTTP Client (Requests)

    The requests library is the backbone for sending GET and POST requests to job board APIs or public pages. It handles cookies, headers, and session persistence with minimal code.

    2. HTML Parser (BeautifulSoup)

    Once the raw HTML is retrieved, BeautifulSoup parses the DOM, allowing you to extract job titles, company names, locations, and application links using CSS selectors or XPath equivalents.

    3. Browser Automation (Selenium or Playwright)

    Many modern career portals rely on JavaScript to render forms or protect against bots. Selenium WebDriver (or the newer Playwright) can simulate real user interactions—clicking buttons, uploading resumes, and solving simple CAPTCHAs.

    4. Data Storage (SQLite, PostgreSQL, or CSV)

    Persisting scraped data enables tracking, analytics, and re‑application logic. For lightweight projects, a CSV file suffices; for larger pipelines, a relational database offers query flexibility.

    5. Scheduler (Cron or APScheduler)

    To keep your job feed fresh, schedule the scraper to run at regular intervals (e.g., every 4 hours). This ensures you never miss a new opportunity.

    Step‑by‑Step Guide to Building Your Own Scraper

    Step 1: Set Up the Project Environment

    python -m venv job-scraper-env
    source job-scraper-env/bin/activate   # On Windows: job-scraper-env\Scripts\activate
    pip install requests beautifulsoup4 selenium pandas
    # Optional: pip install playwright
    

    Step 2: Identify Target Job Boards

    Choose websites that list jobs relevant to your field. Popular options include:

    • Indeed
    • LinkedIn Jobs
    • Glassdoor
    • Company career pages (e.g., https://careers.example.com)

    Check each site’s robots.txt and terms of service to ensure you’re scraping responsibly.

    Step 3: Fetch the Listings

    Below is a minimal example using requests to pull JSON data from a public API endpoint (replace API_URL with the actual URL):

    import requests
    
    API_URL = "https://jobs.github.com/positions.json?description=python"
    response = requests.get(API_URL, headers={"User-Agent": "job-scraper-bot/1.0"})
    response.raise_for_status()
    jobs = response.json()
    print(f"Found {len(jobs)} Python jobs")
    

    Step 4: Parse HTML with BeautifulSoup

    If the site returns HTML, use BeautifulSoup to extract the required fields:

    from bs4 import BeautifulSoup
    
    html = response.text
    soup = BeautifulSoup(html, "html.parser")
    job_cards = soup.select("div.jobsearch-SerpJobCard")
    
    for card in job_cards:
        title = card.select_one("h2.title").get_text(strip=True)
        company = card.select_one("span.company").get_text(strip=True)
        location = card.select_one("div.location").get_text(strip=True)
        link = "https://www.indeed.com" + card.select_one("a")[ "href"]
        print(title, company, location, link)
    

    Step 5: Automate the Application Form

    Many career portals require form submission. Selenium can fill fields and upload files as shown:

    from selenium import webdriver
    from selenium.webdriver.common.by import By
    from selenium.webdriver.common.keys import Keys
    import time
    
    driver = webdriver.Chrome()  # Ensure chromedriver is in PATH
    driver.get("https://jobs.example.com/apply/12345")
    
    # Fill text fields
    driver.find_element(By.NAME, "first_name").send_keys("John")
    driver.find_element(By.NAME, "last_name").send_keys("Doe")
    driver.find_element(By.NAME, "email").send_keys("john.doe@email.com")
    
    # Upload resume
    resume_path = "/path/to/resume.pdf"
    driver.find_element(By.NAME, "resume").send_keys(resume_path)
    
    # Submit the form
    driver.find_element(By.XPATH, "//button[text()='Submit']").click()
    time.sleep(5)  # Wait for confirmation
    driver.quit()
    

    Step 6: Store the Results

    Saving each successful application helps you avoid duplicates and track follow‑ups:

    import pandas as pd
    
    df = pd.DataFrame(jobs)
    df["applied_at"] = pd.Timestamp.now()
    df.to_csv("applied_jobs.csv", mode="a", header=False, index=False)
    

    Step 7: Schedule the Scraper

    Using APScheduler you can run the job every 4 hours without leaving your script:

    from apscheduler.schedulers.blocking import BlockingScheduler
    
    def run_scraper():
        # Call the functions defined above
        fetch_and_apply()
    
    scheduler = BlockingScheduler()
    scheduler.add_job(run_scraper, "interval", hours=4)
    scheduler.start()
    

    Handling Anti‑Scraping Measures Ethically

    Websites employ rate limiting, CAPTCHAs, and IP blocking to protect their data. Ignoring these safeguards can lead to legal trouble or a blocked IP address. Follow these best practices:

    1. Respect robots.txt: If a path is disallowed, avoid scraping it.
    2. Throttle Requests: Insert random delays (2‑5 seconds) between requests.
    3. Rotate User‑Agents: Mimic different browsers to reduce detection.
    4. Use Proxies Wisely: For high‑volume scraping, rotate residential proxies, but always stay within the site’s usage policy.
    5. Handle CAPTCHAs: Prefer services with an API (e.g., 2Captcha) or design your scraper to pause for manual solving.

    Deploying and Scaling Your Scraper

    Containerization with Docker

    Package your scraper in a Docker image to ensure consistent runtime across environments:

    FROM python:3.11-slim
    WORKDIR /app
    COPY requirements.txt .
    RUN pip install -r requirements.txt
    COPY . .
    CMD ["python", "main.py"]
    

    Cloud Scheduling

    Leverage cloud platforms (AWS Lambda, Google Cloud Functions, or Azure Functions) for serverless execution. Pair them with a managed scheduler like AWS EventBridge to trigger the function at desired intervals.

    Monitoring and Alerts

    Integrate logging (via logging module) and alerting (e.g., Slack webhook) to be notified of failures, rate‑limit hits, or successful applications.

    Common Pitfalls and Best Practices

    • Hard‑coding URLs: Use configuration files or environment variables to keep URLs and credentials flexible.
    • Ignoring pagination: Many job boards split results across pages; ensure your scraper iterates through all pages.
    • Storing plain‑text credentials: Secure API keys and login details using secret managers or .env files.
    • Over‑posting: Apply to every scraped job without qualification can hurt your reputation. Implement keyword matching and relevance scoring.
    • Legal compliance: Always review the Terms of Service for each site and consider obtaining permission for large‑scale data extraction.

    Conclusion

    Building

  • Python Dynamic Content Scraping Pyppeteer

    Scraping dynamic, JavaScript‑driven websites is one of the biggest challenges for Python developers today. Traditional tools like requests and BeautifulSoup fetch only the raw HTML that the server returns, leaving you blind to content that is rendered after page load. Pyppeteer—the Python port of Google’s headless‑browser automation library Puppeteer—fills that gap by giving you full control over a Chromium instance, allowing you to wait for scripts, interact with page elements, and extract the data you need. In this guide we’ll explore everything you need to know to master Python dynamic content scraping with Pyppeteer, from installation to advanced anti‑bot techniques, while keeping SEO best practices in mind.

    Why Traditional Scraping Struggles with Dynamic Content

    Modern web applications rely heavily on client‑side JavaScript frameworks such as React, Vue, and Angular. These frameworks populate the DOM after the initial HTML response, meaning:

    • The server‑side HTML often contains only placeholder elements.
    • Critical data (prices, comments, user profiles) is loaded via XHR or fetch calls after the page renders.
    • Content may be hidden behind lazy‑loading, infinite scroll, or interactive tabs.

    When you use requests.get() you receive the pre‑rendered source, so the data you need simply isn’t there. That’s why a headless browser that can execute JavaScript—like Chrome or Chromium—has become the go‑to solution for Python dynamic content scraping.

    Introducing Pyppeteer: The Python Bridge to Headless Chrome

    Pyppeteer mirrors the API of the original Node.js Puppeteer library, providing a clean, async‑first interface to control Chromium. Its key advantages are:

    • Full JavaScript execution: The page behaves exactly as it would in a real browser.
    • Built‑in waiting mechanisms: Wait for selectors, network idle, or custom functions before extracting data.
    • Stealth capabilities: With additional plugins you can mimic human behavior and bypass basic bot detection.
    • Cross‑platform support: Works on Windows, macOS, and Linux without needing a full Chrome installation (it can download a bundled Chromium).

    Setting Up Pyppeteer in Your Python Environment

    Before diving into code, make sure you have a recent version of Python (3.7+ recommended) and pip. Install Pyppeteer with the following command:

    pip install pyppeteer

    Pyppeteer will automatically download a compatible Chromium binary on first run. If you prefer to use an existing Chrome installation, set the executablePath option when launching the browser.

    Core Concepts: Async/Await, Browser, and Page Objects

    Pyppeteer is built around asyncio, so every interaction with the browser is an asynchronous coroutine. Understanding the three main objects will make your code cleaner:

    • Browser: Represents the running Chromium instance. You can open multiple tabs (pages) from a single browser.
    • Page: Analogous to a single tab. All navigation, DOM queries, and screenshot actions happen on a page.
    • Network/Request objects: Allow you to intercept, modify, or block HTTP requests—useful for speeding up scraping or bypassing trackers.

    Basic Example: Scraping a JavaScript‑Rendered Page

    Below is a minimal, fully‑functional script that navigates to a dynamic site, waits for a specific element, and extracts its text content.

    import asyncio
    from pyppeteer import launch
    
    async def scrape():
        # Launch headless Chromium
        browser = await launch(headless=True,
                               args=['--no-sandbox', '--disable-setuid-sandbox'])
        page = await browser.newPage()
    
        # Optional: set a realistic user‑agent
        await page.setUserAgent(
            'Mozilla/5.0 (Windows NT 10.0; Win64; x64) '
            'AppleWebKit/537.36 (KHTML, like Gecko) '
            'Chrome/124.0.0.0 Safari/537.36'
        )
    
        # Navigate to the target URL
        await page.goto('https://example.com/dynamic', waitUntil='networkidle2')
    
        # Wait for the element that holds the data
        await page.waitForSelector('.price-tag')
    
        # Extract the text content
        price = await page.evaluate('''() => {
            return document.querySelector('.price-tag').innerText;
        }''')
    
        print(f'Current price: {price}')
    
        await browser.close()
    
    # Run the async function
    asyncio.get_event_loop().run_until_complete(scrape())

    Key points to notice:

    • waitUntil='networkidle2' tells Chromium to consider navigation complete when there are no more than 2 network connections for at least 500 ms.
    • waitForSelector ensures the script pauses until the target element appears in the DOM.
    • page.evaluate runs a small JavaScript snippet inside the page context, returning the desired value.

    Advanced Techniques for Robust Dynamic Scraping

    1. Waiting for Complex Conditions

    Sometimes a simple selector isn’t enough. Use page.waitForFunction to wait for arbitrary JavaScript conditions:

    await page.waitForFunction(
        'document.querySelectorAll(".item").length > 20',
        timeout=15000
    )

    2. Intercepting and Modifying Requests

    Blocking unnecessary resources (images, ads, analytics) can dramatically speed up scraping and reduce bandwidth. Here’s how to enable request interception:

    await page.setRequestInterception(True)
    
    @page.on('request')
    async def intercept(request):
        if request.resourceType in ['image', 'stylesheet', 'font']:
            await request.abort()
        else:
            await request.continue_()
    

    3. Handling Pagination and Infinite Scroll

    For sites that load more items on scroll, simulate user actions:

    while True:
        # Scroll to the bottom
        await page.evaluate('window.scrollTo(0, document.body.scrollHeight);')
        # Wait for new items to load
        await asyncio.sleep(2)
    
        # Break condition: no new items after a timeout
        new_height = await page.evaluate('document.body.scrollHeight')
        if new_height == last_height:
            break
        last_height = new_height
    

    4. Dealing with CAPTCHAs and Bot Detection

    Many sites employ simple challenges like reCAPTCHA v2 or Cloudflare checks. While Pyppeteer can’t solve CAPTCHAs out of the box, you can:

    • Integrate third‑party solving services (2Captcha, Anti‑Captcha) via their APIs.
    • Use the stealth plugin (available as pyppeteer_stealth) to mask typical automation fingerprints (WebGL, navigator.plugins, etc.).
    • Introduce random delays, mouse movements, and realistic viewport sizes to mimic human browsing.

    Saving Extracted Data: CSV, JSON, and Databases

    After you’ve collected the information, you’ll likely want to store it. Python’s standard libraries make this straightforward.

    Export to CSV

    import csv
    
    with open('prices.csv', 'w', newline='', encoding='utf-8') as f:
        writer = csv.writer(f)
        writer.writerow(['Product', 'Price'])
        for product, price in data:
            writer.writerow([product, price])
    

    Export to JSON

    import json
    
    with open('data.json', 'w', encoding='utf-8') as f:
        json.dump(data, f, ensure_ascii=False, indent=2)
    

    Insert into a Database (SQLite Example)

    import sqlite3
    
    conn = sqlite3.connect('scraped.db')
    c = conn.cursor()
    c.execute('''CREATE TABLE IF NOT EXISTS products
                 (id INTEGER PRIMARY KEY, name TEXT, price TEXT)''')
    c.executemany('INSERT INTO products (name, price) VALUES (?, ?)', data)
    conn.commit()
    conn.close()
    

    SEO‑Friendly Best Practices for Your Scraping Scripts

    Even though your script runs behind the scenes, following SEO‑oriented guidelines can improve discoverability of the content you publish and keep your code maintainable:

    • Use descriptive variable and function names: Search engines index code snippets in tutorials, so clear naming helps both humans and bots.
    • Include relevant keywords naturally: Phrases like “Python dynamic content scraping”, “Pyppeteer tutorial”, and “headless Chrome” should appear in headings, alt‑text (if you embed screenshots), and meta‑descriptions
  • Python Captcha Solver Integration Guide

    Solving CAPTCHAs programmatically can be a game‑changer for developers building web scrapers, automated testing suites, or bots that need to interact with protected forms. In this Python captcha solver integration guide, we’ll walk you through everything you need to know—from choosing the right library, handling different CAPTCHA types, to implementing a robust solution that respects ethical and legal considerations. By the end of this article, you’ll be able to embed a reliable captcha‑solving workflow directly into your Python projects.

    Why Integrate a CAPTCHA Solver in Python?

    CAPTCHAs (Completely Automated Public Turing test to tell Computers and Humans Apart) are designed to block bots, but they can also hinder legitimate automation tasks. Integrating a solver offers several benefits:

    • Increased efficiency: Eliminate manual intervention for repetitive form submissions.
    • Scalability: Automate large‑scale data collection without hitting roadblocks.
    • Improved testing: Simulate real‑user interactions in end‑to‑end test suites.
    • Cost reduction: Avoid hiring human solvers for high‑volume tasks.

    Choosing the Right CAPTCHA Solver

    Not all CAPTCHA solvers are created equal. The best choice depends on the type of CAPTCHA you encounter and your project’s budget.

    1. Open‑Source Libraries

    • pytesseract – Ideal for simple image‑based CAPTCHAs with clear text.
    • captcha_solver – A wrapper around third‑party services; supports reCAPTCHA v2/v3.
    • deathbycaptcha – Provides an API client for the DeathByCaptcha service.

    2. Commercial APIs

    • 2Captcha – Low‑cost, supports image, audio, and Google reCAPTCHA.
    • Anti‑Captcha – Fast response times, offers a Python SDK.
    • CapMonster – High‑throughput solution for enterprise‑scale projects.

    3. Machine‑Learning Approaches

    If you need a fully custom solution, training a deep‑learning model (e.g., using TensorFlow or PyTorch) can give you complete control, but it requires a substantial dataset and compute resources.

    Step‑by‑Step Integration Guide

    Below is a practical, end‑to‑end example that demonstrates how to integrate the popular 2Captcha service into a Python script. The same principles apply to other providers; just swap the API endpoint and request format.

    Prerequisites

    1. Python 3.8+ installed.
    2. A 2Captcha account (free trial available).
    3. Basic knowledge of requests and BeautifulSoup for web scraping.

    1. Install Required Packages

    pip install requests beautifulsoup4

    2. Set Up Configuration

    Store your API key securely—never hard‑code it in public repositories.

    import os
    
    API_KEY = os.getenv('CAPTCHA_API_KEY')  # Export this variable in your environment
    SOLVER_URL = 'http://2captcha.com/in.php'
    RESULT_URL = 'http://2captcha.com/res.php'

    3. Submit the CAPTCHA Image for Solving

    Assume you have already downloaded the CAPTCHA image from the target site and saved it as captcha.png.

    import requests
    
    def submit_captcha(image_path):
        with open(image_path, 'rb') as img:
            files = {'file': img}
            data = {
                'key': API_KEY,
                'method': 'post',
                'json': 1
            }
            response = requests.post(SOLVER_URL, files=files, data=data)
            result = response.json()
            if result.get('status') == 1:
                return result['request']  # This is the CAPTCHA ID
            raise Exception('Captcha submission failed: ' + result.get('request'))

    4. Poll for the Solution

    2Captcha typically needs a few seconds to solve the challenge. Poll the result endpoint until the solution is ready.

    import time
    
    def retrieve_solution(captcha_id, timeout=120, poll_interval=5):
        params = {
            'key': API_KEY,
            'action': 'get',
            'id': captcha_id,
            'json': 1
        }
        elapsed = 0
        while elapsed < timeout:
            response = requests.get(RESULT_URL, params=params)
            result = response.json()
            if result.get('status') == 1:
                return result['request']  # The solved text
            elif result.get('request') != 'CAPCHA_NOT_READY':
                raise Exception('Error retrieving solution: ' + result.get('request'))
            time.sleep(poll_interval)
            elapsed += poll_interval
        raise TimeoutError('Captcha solving timed out')

    5. Submit the Solved CAPTCHA to the Target Site

    Now that you have the solution, include it in the form data and complete the request.

    from bs4 import BeautifulSoup
    
    def submit_form(target_url, solved_captcha):
        session = requests.Session()
        # First, get the page to extract hidden fields (e.g., CSRF tokens)
        page = session.get(target_url)
        soup = BeautifulSoup(page.text, 'html.parser')
        hidden_inputs = {inp['name']: inp.get('value', '') for inp in soup.find_all('input', type='hidden')}
    
        payload = {
            **hidden_inputs,
            'captcha_field_name': solved_captcha,  # Replace with actual field name
            'other_field': 'value'
        }
        response = session.post(target_url, data=payload)
        return response

    6. Full Workflow Example

    def solve_and_submit(target_url, captcha_image_path):
        captcha_id = submit_captcha(captcha_image_path)
        solved_text = retrieve_solution(captcha_id)
        result = submit_form(target_url, solved_text)
        print('Form submitted, status code:', result.status_code)

    Handling Different CAPTCHA Types

    While image CAPTCHAs are the most common, modern sites use more sophisticated challenges. Here’s how to adapt the integration for each type.

    Google reCAPTCHA v2 (“I’m not a robot” Checkbox)

    • Use the sitekey embedded in the page’s HTML.
    • Send a request to the solver with method=userrecaptcha and include googlekey and pageurl.
    • The solver returns a token that you must include in the g-recaptcha-response field when posting the form.

    Google reCAPTCHA v3 (Score‑Based)

    Since v3 runs in the background, you typically need a token generated by the client side. Some services provide a “bypass” endpoint that simulates the JavaScript challenge. Use it cautiously, as many sites monitor for abnormal scores.

    Audio CAPTCHAs

    Audio challenges are easier for OCR engines. Download the audio file, convert it to WAV if needed, and feed it to a speech‑to‑text API (e.g., Google Speech API) or a dedicated audio CAPTCHA solver.

    Best Practices for a Reliable Solver Integration

    • Rate limiting: Respect the provider’s request limits to avoid bans.
    • Error handling: Implement retries with exponential backoff for network glitches.
    • Timeout management: Set reasonable timeouts (e.g., 60‑120 seconds) to prevent hanging scripts.
    • Logging: Record each CAPTCHA ID, solution time, and result for audit trails.
    • Ethical use: Only solve CAPTCHAs for sites where you have permission or for personal testing.

    Testing Your Integration

    Before deploying to production, run the following checks:

    1. Unit tests: Mock the API responses using unittest.mock to verify your logic.
    2. Load tests: Simulate multiple concurrent solves to gauge performance and cost.
    3. Failure scenarios: Force a “CAPTCHA_NOT_READY” response and ensure your polling logic handles it gracefully.

    Common Pitfalls and How to Avoid Them

    • Hard‑coding API keys: Leads to security leaks. Use environment variables or secret managers.
    • Ignoring hidden form fields: Missing CSRF tokens often results in 403 errors.
    • Using the wrong field name: Inspect the page source to locate the exact name attribute for the CAPTCHA response.
    • Overlooking rate limits: Excessive calls can suspend your account; always implement a delay between solves.

    Advanced Topics: Building Your Own Solver with Deep Learning

    If third‑party services don’t meet your latency or cost requirements, consider training a convolutional neural network (CNN) to recognize text in distorted images.

    • Dataset collection: Scrape thousands of labeled CAPTCHA images.
    • Model architecture: Use tf.keras.Sequential with Conv2D, MaxPooling, and Dense layers.
    • Training: Apply data augmentation (rotation, noise) to improve robustness.
    • Deployment: Serve the model via a Flask API and call it from your main script
  • Python Proxy Rotator For Web Scraping

    Web scraping has become an essential tool for data‑driven businesses, but staying undetected is a constant challenge. One of the most effective ways to avoid IP bans and throttling is to rotate proxies automatically. In this guide we’ll walk through everything you need to know about building a Python proxy rotator for web scraping, from choosing the right proxy provider to implementing a robust, reusable solution that scales with your crawling needs.

    Why a Proxy Rotator Is Crucial for Web Scraping

    When you send hundreds or thousands of requests to a target website, the server can easily spot patterns in your traffic—such as a single IP address making repeated calls. Modern anti‑scraping mechanisms use rate limiting, CAPTCHAs, and even machine‑learning models to block suspicious activity. A proxy rotator solves these problems by:

    • Distributing requests across multiple IP addresses, reducing the chance of detection.
    • Bypassing geo‑restrictions, allowing you to scrape region‑specific content.
    • Improving request speed by leveraging parallel connections through different proxies.
    • Providing redundancy: if one proxy fails, the rotator can switch to another without stopping the crawl.

    Choosing the Right Proxy Service

    Not all proxies are created equal. For a reliable rotator you’ll want to consider the following factors:

    Proxy Types

    • HTTP/HTTPS proxies – Ideal for most web scraping tasks that use the requests library.
    • SOCKS5 proxies – Useful when you need to tunnel traffic for protocols beyond HTTP, such as FTP or SMTP.
    • Residential proxies – Provide IPs assigned by ISPs, making them look like real users and reducing block rates.
    • Datacenter proxies – Faster and cheaper, but more likely to be flagged by sophisticated anti‑bot systems.

    Key Features to Look For

    • Rotating endpoint – Some providers give you a single URL that automatically rotates IPs on each request.
    • IP freshness – Fresh IPs reduce the chance of being blacklisted.
    • Authentication method – Basic auth, token‑based, or IP whitelisting.
    • Rate limits per IP – Ensure the provider’s limits align with your scraping speed.

    Core Concepts Behind a Python Proxy Rotator

    Before diving into code, let’s outline the essential components of a robust rotator:

    1. Proxy pool: A collection (list, queue, or database) of proxy URLs.
    2. Health checker: Periodically validates each proxy to keep the pool fresh.
    3. Rotation strategy: Determines how proxies are selected (random, round‑robin, weighted).
    4. Retry logic: Handles failed requests by switching to a new proxy automatically.
    5. Integration layer: Seamlessly plugs into your existing scraping code (e.g., requests or aiohttp).

    Step‑by‑Step Implementation

    1. Setting Up the Environment

    pip install requests beautifulsoup4 lxml

    We’ll use the popular requests library for HTTP calls and BeautifulSoup for parsing HTML. If you prefer asynchronous scraping, replace requests with aiohttp and adjust the code accordingly.

    2. Defining the Proxy Pool

    You can store proxies in a simple Python list, a CSV file, or a Redis set for larger projects. Below is a minimal example using a list of HTTP proxies with basic authentication.

    PROXY_POOL = [
        "http://user1:pass1@192.0.2.1:8080",
        "http://user2:pass2@192.0.2.2:8080",
        "http://user3:pass3@192.0.2.3:8080",
        # Add as many as you need
    ]

    3. Building a Health Checker

    Running a health check before each request ensures you don’t waste time on dead proxies.

    import requests
    from random import choice
    
    def is_proxy_working(proxy_url, test_url="https://httpbin.org/ip", timeout=5):
        try:
            response = requests.get(test_url, proxies={"http": proxy_url, "https": proxy_url}, timeout=timeout)
            return response.status_code == 200
        except Exception:
            return False
    
    def get_live_proxy():
        # Shuffle the pool for randomness
        shuffled = PROXY_POOL[:]
        random.shuffle(shuffled)
        for proxy in shuffled:
            if is_proxy_working(proxy):
                return proxy
        raise RuntimeError("No working proxies found.")

    4. Implementing the Rotator Logic

    The core function wraps any HTTP request, automatically switching proxies on failure.

    def fetch(url, max_retries=5, **kwargs):
        attempt = 0
        while attempt < max_retries:
            proxy = get_live_proxy()
            try:
                response = requests.get(
                    url,
                    proxies={"http": proxy, "https": proxy},
                    timeout=10,
                    **kwargs
                )
                response.raise_for_status()
                return response
            except (requests.RequestException, requests.HTTPError) as e:
                # Log the failure and try the next proxy
                print(f"Proxy {proxy} failed ({e}); rotating...")
                attempt += 1
        raise RuntimeError(f"All {max_retries} attempts failed for {url}")

    5. Parsing the Result with BeautifulSoup

    Now you can combine the rotator with a parser in a clean, reusable way.

    from bs4 import BeautifulSoup
    
    def scrape_product_page(product_url):
        response = fetch(product_url)
        soup = BeautifulSoup(response.text, "lxml")
        title = soup.select_one("h1.product-title").get_text(strip=True)
        price = soup.select_one("span.price").get_text(strip=True)
        return {"title": title, "price": price}

    6. Scaling Up with Threading or AsyncIO

    For large‑scale crawls, you’ll want to run many requests in parallel. Below is a quick example using concurrent.futures.ThreadPoolExecutor:

    from concurrent.futures import ThreadPoolExecutor, as_completed
    
    def batch_scrape(urls):
        results = []
        with ThreadPoolExecutor(max_workers=10) as executor:
            future_to_url = {executor.submit(fetch, url): url for url in urls}
            for future in as_completed(future_to_url):
                url = future_to_url[future]
                try:
                    resp = future.result()
                    results.append({"url": url, "status": resp.status_code})
                except Exception as exc:
                    results.append({"url": url, "error": str(exc)})
        return results

    Best Practices & Tips for a Reliable Rotator

    • Rotate User‑Agents: Pair proxy rotation with a rotating User‑Agent header to mimic different browsers.
    • Respect robots.txt: Even though proxies hide your identity, ethical scraping means honoring a site’s crawling policies.
    • Implement exponential backoff: When a proxy repeatedly fails, wait longer before retrying to avoid hammering the provider.
    • Log every request: Store proxy used, response code, and timestamps for debugging and analytics.
    • Use a persistent storage: For long‑running projects, keep the proxy pool in Redis or a database so health status survives restarts.

    Common Pitfalls and How to Avoid Them

    1. Over‑using free proxies – Free lists often contain dead or blacklisted IPs, leading to high failure rates. Invest in a reputable provider or maintain a strict health‑check routine.

    2. Ignoring SSL verification – Disabling verify=False can expose you to man‑in‑the‑middle attacks. Keep verification on unless you have a compelling reason.

    3. Not handling CAPTCHAs – Some sites serve CAPTCHAs even with rotating proxies. Combine your rotator with a CAPTCHA‑solving service or implement human‑like delays.

    4. Forgetting to rotate cookies – Session cookies can tie requests to a single IP. Clear or rotate cookies when switching proxies.

    Putting It All Together – A Minimal Yet Powerful Scraper

    The following script demonstrates a complete, ready‑to‑run scraper that pulls product titles and prices from a list of URLs while rotating proxies, user‑agents, and handling retries.

    import random
    import requests
    from bs4 import BeautifulSoup
    from concurrent.futures import ThreadPoolExecutor, as_completed

    PROXY_POOL = [
    "http://user:pass@proxy1.example.com:8000",
    "http://user:pass@proxy2.example.com:8000",
    # ... more proxies
    ]

    USER_AGENTS = [
    "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36",
    "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.6 Safari/605.1.15",
    # ... more agents
    ]

    def get_live_proxy():
    random.shuffle(PROXY

  • Python Scrapy Framework Large Scale Crawler

    When it comes to harvesting massive amounts of data from the web, the Python Scrapy framework stands out as a battle‑tested, flexible solution that can power everything from a single‑page scraper to a large‑scale crawler handling millions of requests per day. In this guide we’ll explore how Scrapy’s architecture, built‑in components, and ecosystem extensions enable you to design, optimize, and deploy a crawler that scales horizontally, stays resilient under heavy load, and remains SEO‑friendly for the sites you target.

    Why Scrapy Is the Go‑to Choice for Large‑Scale Crawling

    Before diving into the technical details, let’s recap the key reasons why Scrapy is preferred by enterprises, data‑science teams, and hobbyists alike:

    • Asynchronous networking powered by Twisted, allowing thousands of concurrent requests without blocking.
    • Modular design with spiders, pipelines, and middlewares that can be swapped or extended.
    • Built‑in support for handling cookies, redirects, retries, and auto‑throttling.
    • Extensive ecosystem – Scrapy Cloud, Scrapy Cluster, scrapy‑redis, and many third‑party extensions.
    • Pythonic API that integrates seamlessly with data‑processing libraries like Pandas and SQLAlchemy.

    Core Components of a Scrapy Crawler

    Understanding Scrapy’s building blocks is essential before you start scaling. Each component plays a specific role in the request‑response lifecycle.

    1. Spider

    The spider defines start_urls, parsing logic, and how new requests are generated. For large‑scale jobs you’ll typically write a BaseSpider that other spiders inherit from, centralizing common settings and utilities.

    2. Scheduler & Downloader

    The scheduler queues requests, while the downloader fetches pages. Both are asynchronous, but you can replace the default scheduler with a distributed one (e.g., scrapy_redis.scheduler.Scheduler) to share the queue across multiple machines.

    3. Item Pipeline

    After a spider extracts data into Item objects, pipelines clean, validate, and store the data. For high‑throughput pipelines you’ll want to batch inserts and use asynchronous database drivers.

    4. Middleware

    Middlewares sit between the engine and the downloader/spider, allowing you to modify requests, responses, or handle errors globally. Common uses include rotating proxies, user‑agent rotation, and custom retry logic.

    Scaling Strategies for a Massive Crawl

    Scrapy can run on a single machine for modest workloads, but true large‑scale crawling requires horizontal scaling and robust infrastructure. Below are three proven approaches.

    2.1. Distributed Queues with scrapy-redis

    scrapy-redis replaces the default scheduler and duplicate filter with Redis‑backed versions, enabling multiple Scrapy instances to share the same request queue.

    # settings.py
    SCHEDULER = "scrapy_redis.scheduler.Scheduler"
    DUPEFILTER_CLASS = "scrapy_redis.dupefilter.RFPDupeFilter"
    REDIS_URL = "redis://localhost:6379"
    

    Key benefits:

    • Automatic request de‑duplication across workers.
    • Fault tolerance – if a worker crashes, the queue remains intact.
    • Easy horizontal scaling by adding more Scrapy containers.

    2.2. Scrapy Cluster

    Scrapy Cluster is a full‑featured, container‑oriented architecture built on Kafka, Redis, and Docker. It provides:

    • Message‑driven request distribution via Kafka topics.
    • Stateless workers that can be auto‑scaled in Kubernetes.
    • Centralized monitoring and logging.

    Deploying Scrapy Cluster typically involves three services:

    1. Kafka – holds the request queue and distributes URLs to workers.
    2. Redis – stores duplicate filters, spider stats, and temporary data.
    3. Scrapy Workers – run the actual spiders inside Docker containers.

    2.3. Scrapy Cloud (Portia & Crawlera)

    If you prefer a managed solution, Scrapy Cloud (now part of Zyte) offers:

    • Auto‑scaling infrastructure without manual Docker orchestration.
    • Integrated Crawlera proxy service for IP rotation and anti‑bot evasion.
    • Web UI for spider deployment, logs, and data export.

    While the cost scales with usage, the operational overhead drops dramatically, making it ideal for teams without dedicated DevOps resources.

    Performance Tuning Tips for High‑Throughput Crawls

    Even with a distributed architecture, each Scrapy worker must be fine‑tuned to squeeze maximum performance out of the network and CPU.

    Concurrency Settings

    # settings.py
    CONCURRENT_REQUESTS = 100          # total concurrent requests per worker
    CONCURRENT_REQUESTS_PER_DOMAIN = 20
    CONCURRENT_REQUESTS_PER_IP = 20
    DOWNLOAD_TIMEOUT = 15
    

    Increasing CONCURRENT_REQUESTS raises throughput but can trigger server bans; always combine it with auto‑throttle and respectful DOWNLOAD_DELAY when needed.

    Auto‑Throttle

    # settings.py
    AUTOTHROTTLE_ENABLED = True
    AUTOTHROTTLE_START_DELAY = 2
    AUTOTHROTTLE_MAX_DELAY = 60
    AUTOTHROTTLE_TARGET_CONCURRENCY = 5.0
    

    The auto‑throttle extension dynamically adjusts the request rate based on server response times, helping you stay under the radar while maintaining speed.

    Efficient Item Pipelines

    • Use scrapy-pipelines-redis to batch items into Redis streams before bulk‑loading into a database.
    • Leverage asynchronous drivers (e.g., aiomysql, asyncpg) for non‑blocking DB writes.
    • Compress large payloads with gzip before storing them.

    Proxy & User‑Agent Rotation

    Large crawls inevitably hit anti‑scraping mechanisms. Rotate proxies and user agents at the request level using a custom downloader middleware:

    class RotateProxyMiddleware:
        def __init__(self, proxy_list):
            self.proxies = proxy_list
    
        def process_request(self, request, spider):
            proxy = random.choice(self.proxies)
            request.meta['proxy'] = proxy
    

    Best Practices for a Respectful, SEO‑Friendly Crawl

    Even though you’re building a crawler, you should still respect the target sites’ SEO policies to avoid legal issues and maintain good web etiquette.

    • Read and obey robots.txt – Scrapy does this automatically when ROBOTSTXT_OBEY = True.
    • Throttle aggressively on sites that show signs of overload (high latency, 429 responses).
    • Identify your crawler with a clear User-Agent and provide contact information.
    • Prefer API endpoints over HTML scraping when available – they’re usually more stable and less taxing.

    Real‑World Example: A Distributed Scrapy Spider for E‑Commerce Listings

    The following minimal spider demonstrates how to combine scrapy-redis with a robust parsing routine. It extracts product titles, prices, and stock status from a paginated catalog.

    import scrapy
    from scrapy_redis.spiders import RedisSpider
    
    class ProductSpider(RedisSpider):
        name = "product_spider"
        redis_key = "product:start_urls"
    
        custom_settings = {
            "ITEM_PIPELINES": {
                "myproject.pipelines.MongoPipeline": 300,
            },
            "DOWNLOAD_DELAY": 0.5,
            "AUTOTHROTTLE_ENABLED": True,
        }
    
        def parse(self, response):
            for product in response.css("div.product-item"):
                yield {
                    "title": product.css("h2.title::text").get().strip(),
                    "price": product.css("span.price::text").re_first(r"\d+.\d+"),
                    "in_stock": bool(product.css("span.in-stock")),
                    "url": response.urljoin(product.css("a::attr(href)").get()),
                }
    
            # Follow pagination links
            next_page = response.css("a.next::attr(href)").get()
            if next_page:
                yield response.follow(next_page, callback=self.parse)
    

    To start the