You will get word-level captions burned into your video, timed to the audio

Project details
Captions that land on the word, not on the sentence. Your audio is transcribed at word level with Whisper, so each word appears exactly when it is spoken, and the captions are rendered into the video with ffmpeg rather than handed to you as a file to fight with.
Before delivery I read the transcript against the video and fix the three things auto-captions reliably get wrong: names, numbers and technical terms, and where the lines break. A caption that reads well with the sound off is worth more than one that is merely accurate.
English and Hebrew both work, and Hebrew is rendered right to left properly, with the word order and the highlight both flowing the right way. For any other language, ask me before you pay and I will tell you honestly whether the transcription will be good enough to be worth it.
Before delivery I read the transcript against the video and fix the three things auto-captions reliably get wrong: names, numbers and technical terms, and where the lines break. A caption that reads well with the sound off is worth more than one that is merely accurate.
English and Hebrew both work, and Hebrew is rendered right to left properly, with the word order and the highlight both flowing the right way. For any other language, ask me before you pay and I will tell you honestly whether the transcription will be good enough to be worth it.
Language
English, HebrewWhat's included
| Service Tiers |
Starter
$30
|
Standard
$70
|
Advanced
$150
|
|---|---|---|---|
| Delivery Time | 2 days | 3 days | 5 days |
Number of Revisions | 1 | 2 | 3 |
Number of Minutes | 3 | 15 | 45 |
Embed Subtitles | |||
Transcription | |||
Translation | - | - | - |
SRT File | - |
Optional add-ons
You can add these on the next page.
Additional Revision
+$15
SRT File
+$10Frequently asked questions
About Daniel
AI Video & Data Pipelines That Fail Loudly | Python, FFmpeg
Yavne, Israel - 11:04 pm local time
Most "AI video" work stops at a script. Mine ends at a rendered file. I build the whole chain: transcription, content-aware cutting, 9:16 reframing, burned-in word-level captions, batch rendering. Hours of raw footage become publish-ready clips without an editor ever opening a timeline.
On the LLM side I ship agents that survive real users:
- An AI SMS receptionist on Cloudflare Workers running Llama 3.3 70B, holding full conversations in English and Hebrew and capturing structured leads.
- A consumer app that returns per-item calories and macros from one photo, using Mistral Small 3.1 24B vision on Workers AI, with free and paid tiers.
- A nightly data pipeline indexing 416,817 job postings across 9 hiring platforms, published and running on a schedule.
What I do well:
- Whisper + FFmpeg pipelines: content-aware cutting, 9:16 reframe, animated captions, batch rendering
- Production LLM agents and API integration: Cloudflare Workers AI, Llama, Mistral, Claude, MCP
- Large-scale scraping and data pipelines that keep running after handoff
- Full deploy and ownership: Cloudflare Workers, D1, Pages, iOS and Android release
Microsoft certified: MCSA and Azure.
If you have hours of raw footage, or an LLM prototype that needs to handle real traffic, send it over and I will tell you exactly how I would build it.
Steps for completing your project
After purchasing the project, send requirements so Daniel can start the project.
Delivery time starts when Daniel receives requirements from you.
Daniel works on your project following the steps below.
Revisions may occur after the delivery date.
Transcript checked against the video
Word-level transcription runs first, then I watch it back and correct names, numbers, terms and line breaks. This is the step a one-click tool skips.
Captions burned in, files delivered
The captions are styled, timed and rendered into the video at your aspect ratio. You get the finished file, and on the higher tiers the SRT and VTT as well.

