I will build a scheduled web scraper and data pipeline with clean deduplicated output
About this Gig
Need fresh web data every day or week without re-running a script by hand? I build a scraper that runs by itself on a timer, cleans what it collects, drops repeated items and hands you a file or database that is ready to use.
What you get: a Python scraper (Playwright or Scrapy where the site needs it), automatic runs on your machine or server, validation of every record, a unique key per item so nothing appears twice, retries when a site is slow, and a log that says in plain words if a run failed and why. Output: CSV, XLSX, JSON or SQLite.
Why me: for a client I built a similar pipeline: daily ingestion of public records, duplicate removal, rule-based scoring, automatic runs and crash recovery. One real month run processed 318,962 records in about 34 minutes with 0 errors. The client is confidential.
Not included: sites behind a login you do not own, CAPTCHA bypass, data the site's terms forbid collecting, hosting or proxy costs. I tell you up front if a site is not a fit.
Technology:
Python
•
Excel
•
Scrapy
•
Playwright
•
Pandas
Technique:
Automated
My Portfolio
FAQ
Which websites can you scrape?
Public pages, including JavaScript-heavy ones. I check the site and its terms first. Sites behind a login you do not own, or that forbid scraping, are not a fit.
How are the scheduled runs set up?
With cron (Linux/Mac) or Task Scheduler (Windows) on your machine or server, or a Docker container in Premium. Hosting costs are yours.
What if the website changes later?
Layout changes can break a scraper. Within the revision window I fix what I delivered; later fixes are a new small order. The error log tells you when a run returns nothing.
What does 'pages scraped' mean?
Fiverr asks for a page limit per package. Here it is the number of pages each scheduled run may fetch: 100 in Basic, 500 in Standard, 1,500 in Premium. Larger jobs: message me first.
