Browse categories
Explore
Fiverr Pro
English
$
USD


I build and optimize AI inference backends for apps that need real GPU compute not just a wrapper around the ChatGPT API.
I've deployed and optimized production AI pipelines (image/video processing, face-swap style generative models) on Modal's serverless GPU infrastructure, including:
- GPU/CPU instance configuration (L4/T4) tuned for cost vs. latency
- Cold-start and warmup optimization to cut wasted GPU-seconds
- GPU memory snapshotting to reduce cold-start time
- Pipeline optimization ~75% processing time reduction on a production workload
- Signal handling and reliability fixes for long-running inference jobs
Mobile App Developer
Languages