I will build a whisper transcription API for your product

M
michaelsizonenk
M
michaelsizonenk
Michael S.

About this gig

I have built a streaming voice pipeline end to end: speech to text with Whisper, text to speech with ElevenLabs, running on AWS Bedrock with conversation memory.


Most transcription gigs wrap the API and stop there. The hard parts come after: long files that time out, latency you can hear in a live call, speaker separation, and a bill that grows faster than your usage.


What you get:

- Whisper transcription with timestamps and language detection

- Streaming, so text arrives while the audio is still playing

- Queueing and retries for long or failed files

- Cost logging per minute of audio


My background is telecom. I hold a Master's in Information Networks from Odessa National Academy of Telecommunications and a CCNA, and I worked on an industrial 5G VoIP platform at Siemens with end to end encryption in a team of 45 to 50.


I work from Ukraine on CET.


Tell me your audio volume and your latency target and I will tell you what is realistic before you order.


Get to know Michael S.

Michael S.

CTO and AI Developer, 18 Years: LangChain, RAG, Voice AI, Python Backends

  • FromUkraine
  • Member sinceFeb 2023
  • Avg. response time1 hour
  • Languages

    Russian, Ukrainian, English, German
I'm Michael, founder and CTO at Meduzzen, 18 years in backend and AI engineering. Most AI projects don't fail on the API call. They fail on latency, hallucination and cost once real users arrive. My team and I build the part that holds up: FastAPI backends, LangChain and RAG, Whisper and ElevenLabs voice, voice agent audits, n8n automation and full stack AI SaaS. Earlier, I worked on a 5G VoIP platform at Siemens, IoT security for WeWork, and blockchain threat detection at Cyvers. Send your repo or workflow and I will tell you honestly what it takes.

My Portfolio