I will build a whisper transcription API for your product


About this gig
I have built a streaming voice pipeline end to end: speech to text with Whisper, text to speech with ElevenLabs, running on AWS Bedrock with conversation memory.
Most transcription gigs wrap the API and stop there. The hard parts come after: long files that time out, latency you can hear in a live call, speaker separation, and a bill that grows faster than your usage.
What you get:
- Whisper transcription with timestamps and language detection
- Streaming, so text arrives while the audio is still playing
- Queueing and retries for long or failed files
- Cost logging per minute of audio
My background is telecom. I hold a Master's in Information Networks from Odessa National Academy of Telecommunications and a CCNA, and I worked on an industrial 5G VoIP platform at Siemens with end to end encryption in a team of 45 to 50.
I work from Ukraine on CET.
Tell me your audio volume and your latency target and I will tell you what is realistic before you order.
Get to know Michael S.
CTO and AI Developer, 18 Years: LangChain, RAG, Voice AI, Python Backends
- FromUkraine
- Member sinceFeb 2023
- Avg. response time1 hour
Languages
Russian, Ukrainian, English, German
My Portfolio
FAQ
Which speech-to-text engine do you use?
Whisper by default, since it handles accents and background noise well. I also work with alternatives like Deepgram or AssemblyAI if your project already runs on one of them, or needs a feature they cover better.
Do you support real-time or streaming transcription?
Yes, from the Standard package. Text streams in while the audio is still playing, instead of you waiting for the whole file to finish before you see anything.
Which languages are supported?
Whisper covers dozens of languages out of the box, including automatic language detection. Send me your target languages and I will confirm coverage and flag any accuracy trade-offs before you order.
Can you add text-to-speech, so my app can talk back?
Yes, on the Premium package. I pair the transcription with ElevenLabs for natural-sounding voice output, so you get a full speech-in, speech-out pipeline, not just transcription.
How do you handle long or failed audio files?
Long files go into a queue with retries from the Standard package up, so one slow or failed job does not block the rest of your uploads or time out the request.
Can this run on my own server or cloud account?
Yes. On the Premium package I deploy the pipeline into your own AWS or Azure environment, so the audio and the infrastructure stay under your control.
How is the cost calculated, and can I keep it under control?
I log cost per minute of audio processed on every package, so you can see exactly what each request costs instead of getting surprised by the monthly bill.
