I will build a data cleaning and validation pipeline in python

N
natsuki_yunyuan
N
natsuki_yunyuan
natsuki

About this gig

Most data problems are not analysis problems. They are cleaning problems nobody wants to look at.


A pipeline that silently "fixes" bad data is worse than one that refuses it. It turns 15/01/2026 into the wrong month, drops 3% of rows, and produces a report that looks clean and is wrong.


WHAT YOU GET

- A Python pipeline that reads your files, cleans what it can prove, and quarantines the rest with a written reason.

- Every rejected row goes to its own file with its source and line number. Nothing is dropped quietly.

- Validation rules you can read and change: types, ranges, required fields, date formats, duplicates.

- An auditable report plus exit codes, so it fails loudly instead of writing a wrong file.

- Tests, so you can change the rules without breaking them.

- A handover doc: what was built, how to run it, what was verified, and the limitations.


HOW I WORK

I look at your data before quoting. If a column is ambiguous I will ask rather than guess.


STACK

Python, standard library only. Nothing to install.


Tell me what "clean" means for your data and I will say if it is a fit. Samples on GitHub - message for the link.

Get to know natsuki

natsuki

all

  • FromChina
  • Member sinceSep 2026
  • Avg. response time1 hour
  • Languages

    Chinese, English
I build automation that connects AI to the real world. AI agents with real memory architecture - not a prompt around a chat API. Data pipelines that refuse to guess. Tools that pull structured data out of human-readable documents. I ship with verification: my samples carry a test suite you can run yourself, one command, no install. I also write handover docs - how to run it, what was verified, and what it does NOT do. Three public samples on GitHub, 99 passing tests - message me for the link. I also work on motion control: EtherCAT, ODrive, Klipper.