Python Web Broken Link Checker Utility

Written by

in

Keeping a website free of dead or misdirected URLs is essential for user experience, SEO rankings, and overall site credibility. A Python‑based broken link checker utility can automate the tedious task of crawling pages, validating hyperlinks, and reporting issues before they hurt your traffic. In this guide we’ll explore why link checking matters, outline the key features of a robust utility, walk through a step‑by‑step implementation, and share best‑practice tips for integrating the tool into your development workflow.

Why a Broken Link Checker Is a Must‑Have for Modern Websites

  • Improves SEO: Search engines downgrade pages with many broken links, which can lower rankings.
  • Boosts user trust: Visitors quickly lose confidence when they encounter 404 errors.
  • Reduces bounce rate: A smooth navigation path keeps users engaged longer.
  • Supports accessibility: Screen readers announce broken links, creating a confusing experience for users with disabilities.
  • Automates maintenance: Manual checking is impractical for large sites; a script can run nightly or on every deployment.

Core Features of an Effective Python Broken Link Checker

  • Recursive crawling of internal pages with a configurable depth limit.
  • Support for http, https, and protocol‑relative URLs.
  • Parallel request handling (threading or async) for speed.
  • Detection of common HTTP status codes (404, 500, 301/302 redirects, 429 throttling).
  • Whitelist/ignore patterns to skip external resources or known false positives.
  • Exportable reports in JSON, CSV, or plain‑text formats.
  • Easy integration with CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins).

Building a Simple Yet Powerful Checker from Scratch

1. Choose the Right Libraries

For a balance of simplicity and performance, combine requests (or httpx for async) with beautifulsoup4 for HTML parsing. If you need high concurrency, aiohttp plus asyncio is a solid choice.

2. Project Structure

broken_link_checker/
│
├─ checker.py          # Core crawling and validation logic
├─ utils.py            # Helper functions (URL normalization, logging)
├─ requirements.txt    # Dependencies
└─ reports/
   └─ latest_report.json

3. Sample Implementation (Synchronous Version)

import re
import json
import queue
import threading
from urllib.parse import urljoin, urlparse

import requests
from bs4 import BeautifulSoup

# ---------- Configuration ----------
START_URL = "https://example.com"
MAX_DEPTH = 3
TIMEOUT = 10
USER_AGENT = "PythonLinkChecker/1.0"
IGNORE_PATTERNS = [r"^mailto:", r"^tel:", r"\.pdf$"]
# ----------------------------------

session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT})

def should_ignore(link):
    return any(re.search(pat, link) for pat in IGNORE_PATTERNS)

def is_internal(link):
    return urlparse(link).netloc in ("", urlparse(START_URL).netloc)

def fetch(url):
    try:
        resp = session.get(url, timeout=TIMEOUT, allow_redirects=True)
        return resp.status_code, resp.url
    except requests.RequestException as e:
        return None, str(e)

def extract_links(html, base):
    soup = BeautifulSoup(html, "html.parser")
    links = set()
    for tag in soup.find_all("a", href=True):
        href = tag["href"].strip()
        if should_ignore(href):
            continue
        absolute = urljoin(base, href)
        links.add(absolute.split("#")[0])  # Drop fragment
    return links

def worker(q, results, visited):
    while True:
        try:
            url, depth = q.get_nowait()
        except queue.Empty:
            break

        if url in visited or depth > MAX_DEPTH:
            q.task_done()
            continue

        visited.add(url)
        status, final_url = fetch(url)
        results.append({"url": url, "status": status, "final_url": final_url})

        if status and 200 <= status < 400 and is_internal(final_url):
            try:
                resp = session.get(final_url, timeout=TIMEOUT)
                child_links = extract_links(resp.text, final_url)
                for link in child_links:
                    q.put((link, depth + 1))
            except requests.RequestException:
                pass

        q.task_done()

def run_checker():
    q = queue.Queue()
    q.put((START_URL, 0))
    results = []
    visited = set()
    threads = []

    for _ in range(8):  # Number of worker threads
        t = threading.Thread(target=worker, args=(q, results, visited))
        t.start()
        threads.append(t)

    q.join()
    for t in threads:
        t.join()

    return results

if __name__ == "__main__":
    report = run_checker()
    with open("reports/latest_report.json", "w", encoding="utf-8") as f:
        json.dump(report, f, indent=2)
    print(f"✅ Scan complete – {len(report)} URLs processed.")

4. Key Points in the Code

  • Queue‑based crawling: Guarantees breadth‑first traversal and easy depth control.
  • Thread pool: Eight workers provide a good speed‑up without overwhelming the target server.
  • Ignore patterns: Regular expressions prevent false positives on mailto links, PDFs, or JavaScript URLs.
  • Result recording: Each entry stores the original URL, HTTP status, and final redirected URL.

Scaling Up with AsyncIO for Massive Sites

If you need to scan thousands of pages, switch to aiohttp and asyncio.Semaphore to limit concurrent connections. Below is a concise async skeleton you can drop into checker.py:

import asyncio
import aiohttp
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlparse

SEM = asyncio.Semaphore(20)  # Max 20 concurrent requests

async def fetch(session, url):
    async with SEM:
        try:
            async with session.get(url, timeout=TIMEOUT) as resp:
                text = await resp.text()
                return resp.status, str(resp.url), text
        except Exception as e:
            return None, str(e), ""

async def crawl(url, depth, session, visited, results):
    if url in visited or depth > MAX_DEPTH:
        return
    visited.add(url)

    status, final_url, html = await fetch(session, url)
    results.append({"url": url, "status": status, "final_url": final_url})

    if status and 200 <= status < 400 and is_internal(final_url):
        soup = BeautifulSoup(html, "html.parser")
        for tag in soup.find_all("a", href=True):
            link = urljoin(final_url, tag["href"]).split("#")[0]
            if not should_ignore(link):
                await crawl(link, depth + 1, session, visited, results)

async def main():
    async with aiohttp.ClientSession(headers={"User-Agent": USER_AGENT}) as session:
        visited = set()
        results = []
        await crawl(START_URL, 0, session, visited, results)
        # Save JSON report...
        ...

if __name__ == "__main__":
    asyncio.run(main())

Integrating the Checker Into Your CI/CD Pipeline

Automated link validation should be part of every deployment. Here’s a quick checklist for a GitHub Actions workflow:

  1. Create a requirements.txt with requests, beautifulsoup4, and optionally aiohttp.
  2. Add a job that installs dependencies and runs python checker.py.
  3. Fail the build if the JSON report contains any entry with status >= 400.
  4. Upload the report as an artifact for later review.
name: Link Check
on: [push, pull_request]

jobs:
  broken-links:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v3
      - name: Set up Python
        uses: actions/setup-python@v4
        with:
          python-version: "3.12"
      - name: Install deps
        run: pip install -r requirements.txt
      - name: Run checker
        run: python checker.py
      - name: Fail on broken links
        run: |
          python -c "
import json, sys
with open('reports/latest_report.json') as f:
    data = json.load(f)
bad = [r for r in data if r['status'] and r['status'] >= 400]
if bad:
    print('Broken links found:', len(bad))
    sys.exit(1)
"
      - name: Upload report
        uses: actions/upload-artifact@v3
        with:
          name: link-report
          path: reports/latest_report.json

Advanced Tips for a Production‑Ready Utility

Handle Rate Limiting Gracefully

  • Detect 429 Too Many Requests and respect the Retry-After header.
  • Implement exponential back‑off for repeated failures.

Cache Results to Speed Up Re‑scans

Store the last known status of each URL in a lightweight SQLite database. On subsequent runs, skip URLs that haven’t changed

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *