I will extract data from pdf and scanned documents into excel
Web scraping, data cleaning and Telegram bots in Python
About this Gig
I turn PDFs, scans and Word files into a clean spreadsheet you can actually use.
Send me the documents and the fields you need - invoices, reports, forms, price lists, registries - and you get a structured Excel or CSV, one row per record. No retyping, no copy-paste, no half-empty columns.
What you get:
- Data in Excel, CSV, JSON or Google Sheets, your choice
- - One row per record, columns named the way you ask
- - A confidence flag on any uncertain field, so you know exactly what to double-check
- - For Premium: OCR for scans, plus a reusable pipeline and a short README
How I work. I look at a few of your real documents before quoting. A clean text PDF and a crooked scan are very different jobs, and I would rather tell you the honest price and accuracy up front than surprise you later.
What I don't do. I don't guess. If a scan is too poor to read a number reliably, I flag it instead of inventing a value, so your totals stay trustworthy.
Send me a sample and I'll tell you exactly what's possible.
Industry:
Financial services & business
•
Manufacturing & storage
Tool:
Excel
•
Google Sheets
•
PDF editor
My Portfolio
FAQ
My scans are low quality or handwritten - can you still do it?
Printed text, even from a crooked scan, I handle with OCR and deskewing. Pure handwriting has no reliable accuracy, so I tell you honestly before you order.
Can I see a sample before I order?
Yes. Send a few documents and I will process 5-10 of them and send the result back before you commit.
Do I get the source code?
For the Premium pipeline, yes, with a short README so you can run it on new batches yourself.
