Choose a file the browser can decode
Start with a recording you are permitted to process. MP3 is a common audio choice; PCM WAV avoids another lossy audio encode. MP4 and M4A commonly contain AAC audio, while WebM often contains Opus. The extension identifies a container, not a promise that every track inside it is supported. A video can play visually and still have an audio codec the transcription browser cannot decode.
Keep the original when troubleshooting. If the local file check reports an unsupported track, try a current desktop Chrome browser first. If a conversion is necessary, use a trusted local converter and retain the original duration and speech. Renaming a filename from MOV to MP4 does not convert its contents. This website does not upload your file for server conversion, and it does not include a universal FFmpeg fallback.
Open the file and choose the spoken language
On the homepage, use Choose File or drag the recording into the upload area. Wait for the audio-track check before starting. Confirm that the displayed filename and duration match the recording you intended to use. Website language controls the interface; it does not set the language spoken in your file. Choose English, French or Arabic separately in the workspace, or use Auto Detect when the language is unknown.
Automatic detection examines the beginning of the recording. If an introduction uses a different language, explicit selection can give a more appropriate starting point. The task is transcription in the spoken language, not translation. The English and French path offers general model choices; a larger choice may cost more time and memory. Arabic follows its dedicated model path regardless of those general model settings. Do not assume a GPU setting accelerates every language.
Let the required model load
Start Transcription loads the selected recognition model when necessary, then processes the audio locally. The first run includes downloading model files, so distinguish model loading from recognition progress. Arabic weights are about 1.33 GB; other model choices have their own download requirements. A browser cache can help a later visit, but it is not permanent storage or a guarantee of offline operation.
Keep the page open and give the device enough available memory. Long recordings take more work even though audio preparation uses windows rather than one enormous decoded buffer. Avoid running multiple transcription tabs on a constrained laptop. Cancel stops the current task; if you need to retry, review the status and restart from the available controls. Save any result you need before refreshing or navigating away. Download speed and processing speed are separate limits, so a faster connection will not necessarily speed up recognition after loading.
Review the transcript against the audio
Once text appears, listen again to important passages in your media player. Check names, dates, amounts, technical vocabulary and negative statements before polishing punctuation. Noise and overlapping speakers may produce confident-looking errors. A readable sentence is not evidence that every word was recognized correctly. For French, pay special attention to names, agreement and homophones; for English, check contractions and repeated words.
Edit the working transcript directly, or open Check & Improve to inspect local suggestions. English uses Harper and French uses Grammalecte; Arabic has limited rule-based cleanup. Suggestions do not replace listening, and a grammar correction can be inappropriate for a quotation or a deliberate speaking style. Accept or ignore each issue according to the recording. The original document remains available if you need to compare or reset. There is no sentence rewriting or summarization step.
Choose the export that matches your purpose
Use Copy Text or TXT for the selected readable document version. If you accepted corrections or edited the working copy, select that version before exporting. SRT and VTT serve subtitle workflows; JSON preserves structured segments for software that can consume them. These timed exports intentionally retain the original recognition text and timestamps. Editing the readable transcript does not automatically correct the subtitle files.
Load SRT or VTT beside the video in your destination player or editor and check synchronization, cue length and readable line breaks. A downloaded subtitle file does not burn text into the video. Segments with missing or invalid timestamps require attention and are not silently converted into invented timed cues. Keep both a reviewed text document and the original timed output when you need an audit trail. Download before closing the page, and share only material you are authorized to distribute. The related tools below all lead to the same local workspace; pick the guide or format that fits your recording.