Skip to content
CutConvert
Post-production file converter

Convert JSON to SRT subtitles

Turn timed transcripts or YouTube chat replay JSON/JSONL into SRT. Preview the detected captions, or map text and timestamp fields in a custom JSON export.

Convert your files

Drop files here

browseWhisper or transcript JSONConverted in your browser. Files never leave your device.About an hour of footage is free. Larger files: $3 each, or Pro $9/month for everything up to 50 MB.

up to 2 files · 2 MB each

free to start · project to project: $5 · every other conversion: $3 · Pro $9/month · pricing

Converting as a guest. A free account has no monthly cap on files within the free size limit.

Queue
0/2
No files queued yet — drop something above to begin.

How to convert JSON to SRT

Drop a supported timed JSON transcript to preview and download SRT subtitles. Whisper exports with timestamped segments work directly. For an unfamiliar JSON layout, map the array, text and timestamp fields under Delivery settings. A transcript without timestamps cannot supply original timing; encrypted data needs a supported export from its source app.

  1. Add your JSON files

    Drag your Whisper or transcript .json files onto the drop zone, or click to browse.

  2. Convert JSON to SRT

    Press Convert. Each timed segment becomes an SRT cue with millisecond timing; speaker fields become labels.

  3. Download your result

    Download the converted file instantly, or a single ZIP archive when you convert a batch of files.

Try JSON to SRT with a sample

See the source file and its converted result before using your own files. Use “Try a sample” in the converter to preview the same input.

See an excerpt of the sample input
{
  "task": "transcribe",
  "language": "english",
  "text": "Welcome back to the cutting room. Today we're converting subtitles in seconds.",
  "segments": [
    { "id": 0, "start": 1.0, "end": 4.0, "text": " Welcome back to the cutting room.", "avg_logprob": -0.2, "no_speech_prob": 0.01 },
    { "id": 1, "start": 4.2, "end": 7.5, "text": " Today we're converting subtitles in seconds.", "avg_logprob": -0.19, "no_speech_prob": 0.01 },
    { "id": 2, "start": 7.8, "end": 11.4, "text": " Drop a whole batch and get one tidy ZIP back.", "avg_logprob": -0.21, "no_speech_prob": 0.02 }
  ]
}
  1. Upload timed JSON and check the detected dialect in the preview.
  2. For an unfamiliar schema, enable Custom JSON fields and specify the array, text, start, end and time units.
  3. Download SRT and import it as captions in your editor; check timing against the source.

Missing timestamps and encrypted drafts cannot be recovered by mapping field names. Re-export timed captions from the source app.

Synthetic examples. Downloaded outputs are checked against the conversion engine; this does not certify an import in your editor version.

Back to the converter

Transcript JSON and SRT subtitles, explained

AI transcription outputs JSON; players and editors want SRT. This converter accepts the shapes that actually exist in the wild: OpenAI Whisper verbose_json (segments with start/end in seconds), whisper.cpp (transcription arrays with millisecond offsets), and generic arrays of {start, end, text} segments. Transcript timing is read from the file. Chat replays retain message start offsets and use estimated display durations, with a warning for review.

Turn timed transcripts or YouTube chat replay JSON/JSONL into SRT. Preview the detected captions, or map text and timestamp fields in a custom JSON export. It is free to start, encrypted in transit, and converts a whole batch into one ZIP — sign in when you need large, high-volume jobs.

JSON to SRT FAQ

Can I convert YouTube live-chat replay JSON?

Yes. Upload the .live_chat.json or JSONL file from yt-dlp. Messages use their video replay offsets and retain author names. Display durations are estimated up to six seconds; simultaneous messages can overlap. This creates chat captions, not a transcript of the spoken audio. Live captures may need timecode alignment.

Can I turn a Twitch chat log into subtitles?

Yes. Drop the chat JSON that TwitchDownloader writes (Chat Download, JSON format). Each message becomes a subtitle line with the commenter's name, messages that arrive within two seconds share one cue of up to three lines, and a cue stays on screen until the next one or six seconds. The result burns in as a chat replay with any subtitle-capable editor or player.

Which JSON formats are supported?

OpenAI Whisper verbose_json (a segments array with start/end in seconds), whisper.cpp output (a transcription array with offsets or timestamps), and generic arrays of {start, end, text} objects, with optional speaker fields, in seconds or milliseconds.

How do I get verbose_json out of Whisper?

With the OpenAI API, set response_format to verbose_json. Running Whisper locally, the default .json output works; whisper.cpp's --output-json file works too.

My JSON only has a plain text field. Can it be converted?

No. SRT needs per-segment start and end times, and a bare text transcript has none. Re-export with segment timing (verbose_json for Whisper) and it will convert.

Are speaker labels from diarization kept?

Yes. Segments carrying a speaker field come out as labeled SRT cues, so diarized transcripts stay attributed.

What if my file is a Premiere Pro transcript JSON?

Use the Premiere Transcript to SRT converter at /prtranscript/srt. It reads Premiere's transcript format (including binary .prtranscript files) directly.

Guides for this workflow

Why editors trust it

Free to start

Convert everyday batches for free. Create an account when you need large, high-volume jobs.

Batch in, one ZIP out

Convert a batch within your allowance and download a single ZIP, or one file if that is all you need.

Made for editors

Whole timelines from Premiere, Final Cut and Avid with transitions, markers, titles, keyframes, levels, nests and bins; captions that keep their italics, colour, placement and speakers across SRT, VTT, ASS, TTML, iTT, SCC and STL, any of them in, any of them out; transcripts into Premiere with the words the engine timed. Frame rates and timecode handled the way post expects.

Private by design

Files are converted in your browser. Queued files can be saved locally to restore your job after checkout. Conversion metadata helps us diagnose errors; file contents are not uploaded.

Questions, answered

Where are my files processed?

In your browser. The converter runs on your own machine and the file never leaves it; the site only records that a conversion happened. CutConvert never sees, stores, sells or shares your media.

Is it free? Do I need an account?

About an hour of footage per file is free, on every converter: 100 KB of subtitles or captions, 500 KB of Word transcript, 2 MB of transcript JSON, 1 MB of Resolve project, 5 MB of Final Cut XML, with Premiere projects and Avid AAF a little under the hour. Guests get five conversions a month, a free account has no monthly cap within those sizes, and above that project-to-project transfers are $5 per file. Every other conversion is $3, including project files to captions, reports or cut lists, and captions or transcripts into an editor. Pro is $9 a month ($90 a year) for every file type up to 50 MB; Studio is $19 a month ($190 a year) for files up to 500 MB. Paid-file access lasts 24 hours.

What do I get when I convert multiple files?

Drop several files and you get back one ZIP archive containing every converted file. Convert a single file and you get that one file back, no ZIP. Tick extra formats under the output chips and the same ZIP carries every one of them, and a file counts as one conversion whatever it is written to.

Which formats are supported?

Timelines: Premiere .prproj, DaVinci Resolve .drp (as exported, or gzip- or Zstandard-wrapped), Avid AAF and Final Cut FCPXML in; FCP7 XML, FCPXML, Resolve .drp, CMX 3600 EDL and shot-list CSV out. A Premiere or Final Cut edit travels with its transitions, markers, titles, opacity and scale keyframes, clip volume, speed changes, nested sequences, bins, audio channel routing and each file's own frame rate, wherever the output has a place for them; an Avid AAF brings its transitions, markers, clip levels with their keyframes, muted tracks, nested stacks and each clip's own rate. Captions and transcripts, read and written: SRT, WebVTT, SBV, ASS, TTML, DFXP, iTT, SCC, MCC (a second 608 channel comes out as its own file), EBU-STL, Avid caption TXT, CapCut drafts of any size, timed CSV, Word, plain text, Whisper, WhisperX, AssemblyAI, Deepgram, Rev.ai and CapCut JSON, TwitchDownloader chat logs as a chat replay, Premiere transcript JSON and .prtranscript, Rev, Otter and Word-style transcripts, with italics, colour, placement and speakers kept where the output can hold them and measured word timing into Premiere. EDLs become event CSV reports. The capability table lists what every pair keeps.

Can I pick a frame rate?

Yes. For EDL and Avid caption output you can choose 23.976, 24, 25, 29.97, or 30 fps, or let CutConvert auto-detect it from your source.