나는 깨지기를 거부하는 자기 치유 웹 스크레이퍼를 만들었습니다.

작성자

카테고리:

← 피드로
DEV Community · Dharmendra Kushwaha · 2026-08-12 개발(SW)

Dharmendra Kushwaha

Every developer who has worked with web scraping knows the pain: a website updates its UI, a CSS class name changes, and suddenly your entire data pipeline breaks.
To solve this, I’ve been building a Self-Healing Web Scraper in Python. Instead of failing silently when a website redesigns, it automatically adapts, heals itself, and continues extracting data.
–> The Problem
“Normal scrapers are incredibly fragile. They rely on rigid HTML structures and CSS selectors. The moment a website changes its layout, your scraper dies”.
The Solution: A 5-Layer Fallback Pipeline
I designed an extraction engine that tries five different strategies in sequence. If one fails, it automatically falls back to the next:
1) JSON-LD Parsing: Looks for structured data (the most resilient method).
2) CSS Selectors: Traditional fast extraction if classes haven’t changed.
3) Regex Fallback: Pattern matching on raw HTML when CSS fails.
4) Fuzzy Self-Healing: Scores nearby text blocks to find price symbols and product data as a last resort.
5) LLM / AI Self-Healing: The newest addition! If all standard methods fail, it uses an LLM (Google Gemini or offline Ollama) to “read” the broken HTML and extract the products anyway.

  • Self-Learning Memory
    The scraper doesn’t just survive; it learns. It remembers which extraction strategy worked best for a specific website and automatically prioritizes that method on the next run, saving execution time and API calls.
    |–| Key Features

    Multi-Site Price Comparison: Tested across Flipkart, Amazon, Robu.in, and more.
    AI-Powered Resilience: Uses Google Gemini or 100% offline Ollama for intelligent data extraction.
    Smart Garbage Filtering: Rejects junk data like menu items, ads, and accessories.
    Auto Currency Conversion: Automatically normalizes prices to ₹ INR.
    Dual Fetching & Stealth: Uses Playwright with stealth mode to bypass basic anti-bot detection.
    Persistent Storage: Saves data to an SQLite database and exports clean, timestamped CSV files.

–>Watch it in Action
I recorded a quick demo showing how the scraper handles failures and triggers the AI fallback:
–>Watch the Demo Video Here
–> Check out the Code
I am open-sourcing this project and would love to get feedback from the community on architecture, edge cases, or feature requests!
🔗 GitHub Repository: https://github.com/uyg7x/Self-Healing-Web-Scraper
If you find this useful or interesting, I would highly appreciate a ⭐ on the repo and your thoughts in the comments below! How do you currently handle website redesigns in your scraping pipelines?

원문에서 계속 ↗

추출 본문 · 출처: dev.to · https://dev.to/pjyprovider/self-healing-scraper-1o21

코멘트

답글 남기기