I will build a website content inventory with urls and metadata
Web Data Collection and Processing
About this Gig
I will turn a buyer-supplied list of publicly accessible website URLs into a clean content inventory.
Your delivery can include the page URL, HTTP status, title, meta description, canonical URL, and duplicate-title or duplicate-description flags. I organize and check the inventory before delivery in CSV, Excel XLSX, or JSON.
For multi-page orders, please provide the exact public page URLs in a file; whole-site page discovery from one homepage is not included.
I only work with public, customer-provided, licensed, or otherwise authorized sources. Login-protected pages, passwords, CAPTCHA bypasses, private systems, and prohibited personal-data collection are not included.
Send a few representative URLs before ordering if you are unsure whether the source is compatible.
Technology:
Python
•
Excel
•
Scrapy
•
Beautiful soup
•
Pandas
Technique:
Automated
My Portfolio
FAQ
Can you discover every page from one homepage?
Not in this package. For multi-page work, please provide the exact public page URLs in a file. I can process one supplied homepage as a one-page inventory.
What fields can the inventory include?
Typical fields are page URL, HTTP status, title, meta description, canonical URL, and duplicate-title or duplicate-description flags.
Do you process private or login-protected pages?
No. I work only with publicly accessible, customer-provided, licensed, or otherwise authorized pages and do not bypass access controls.
What formats can you deliver?
I can deliver CSV, Excel XLSX, or JSON. CSV is usually the simplest choice for a content inventory.
What if some supplied URLs fail?
Failed or unavailable public URLs are retained with their status when possible so the inventory clearly shows what was and was not accessible.
