You can find spoken details in a saved video without replaying the whole recording. OGAD (Off Grid AI Desktop) Pro on Mac converts the audio into a transcript in its Voice library. Search for a word, open the matching recording, and copy the text. The speech model processes the file locally.
| OGAD Pro | Desktop downloads and platform details |

What would you like to do with Off Grid AI?
Have a feature or use case you would like us to support? Tell us what you want to do and which device you use.
Write to support@offgridmobileai.co, join our Slack community, or talk to us on Reddit.
This workflow uses Voice > Transcribe file, a Pro feature. It works with a saved video on your Mac. You do not need to play the video through your speakers and record it again.
For example, import a recorded product demonstration, then search for the name of a feature discussed in it. The transcript helps you find the recording and read its words. You can replay the video to check what was said.
What do you need before you import the video?
You need OGAD Pro on a supported Apple Silicon Mac, an active Pro license, and a downloaded speech model. You also need the video saved locally and enough storage for a copy plus working files. Complete app setup, license activation, and model downloads while you have internet access.
This guide uses the Mac Pro workflow. The desktop platform page currently lists Pro as available on macOS.
Check the video before you import it:
- Confirm that it plays and contains audible speech.
- Make sure the file is downloaded, rather than a cloud-storage placeholder.
- Give it a name you will recognize in the Voice library.
- Keep the original file until you have checked the transcript.
Supported video extensions include MP4, MOV, MKV, AVI, M4V, WebM, 3GP, FLV, and WMV. An accepted extension does not guarantee that every codec inside that container will decode. Start with a short video that plays correctly on your Mac.
Which transcription model should you select?
In Voice settings, set Engine to Whisper, then select a multilingual model for recordings in languages other than English. Whisper Base is a small starting point. Compare it with Whisper Small if the first transcript misses words. Larger models need more storage and available memory.
| Model | Approximate desktop download | When to try it |
|---|---|---|
| Whisper Base | 148 MB | A short first video to check the workflow |
| Whisper Small | 488 MB | A comparison when Base misses speech |
| Whisper Medium | 1.53 GB | A larger option if the Mac has enough memory |
These are rounded decimal file sizes, not RAM requirements. The app can display rounded sizes differently.
The model name appears in the settings panel even though it is not something you need to install through a terminal. For multilingual recordings, avoid English-only files with .en in the filename. The Whisper documentation explains that distinction.
How do you turn a video into a transcript?
Open Voice, select Transcribe file, and choose the saved video. The app adds a recording card and transcribes its audio locally. When processing finishes, the text appears with the recording. Longer transcripts have a Show more control, and Copy transcript copies the text.
1. Set up the speech model
Open Voice, then use the gear button marked Voice settings. Under Transcription, set Engine to Whisper.
Download the model you want from the list in that panel. If it is already downloaded, select Use. Confirm that it shows Active before importing the video.
2. Import the saved video
Return to the Voice library and select Transcribe file. Choose your local video in the file picker.
A new card appears for the file. While it runs, the card shows Transcribing on-device. The app copies the imported file into its own recording storage, so keep enough free disk space for both copies.
You can also drop a file onto the Voice screen. Drag-and-drop imports have a displayed 1 GB size limit. For a first test, use a short file well below that limit.
3. Read and check the transcript
When the text appears, select Show more if it is collapsed. Review a short passage against the video.
Check names, dates, and amounts before using the transcript in a document. Clear formatting does not establish that the speech recognition was correct.
Use the video controls on the recording card to replay the source when playback is available. A codec that can be decoded for transcription may not necessarily play in the embedded player.
4. Copy the text
Use Copy transcript on the recording card. Paste the result into your notes or a document if you want to make corrections or keep a separate copy.
The library provides selectable text and a copy action. These steps do not depend on an in-app transcript editor or a subtitle export feature.
How do you search the transcript later?
Type a word or phrase into Search transcripts on the Voice screen. The filter matches text in the recording title and transcript, without matching letter case. Search with a distinctive word you know appears in the transcript, then read the matching card.
For a first check, find a word in the completed transcript and enter it in the search field. This separates a search problem from a transcription error.
If a person’s name was transcribed incorrectly, searching for the correct spelling will not find that name. Try another nearby word or part of the filename.
There are two limits to understand:
- Search matches the entered text. It is not a meaning-based question-answering search.
- The current Voice screen loads the newest 500 recordings. Its search filters that loaded set, even if your retention setting keeps more recordings.
A matching card does not promise a jump to the exact spoken second. Use the transcript and playback controls to check the passage.
How do you keep the recording available?
Open Voice settings > History and check Keep and Auto-delete. These settings control how many recordings remain and whether older recordings are removed. Set them to match the period for which you need the video and transcript.
Retention can remove the app’s stored media and its transcript entry. It does not replace a backup plan for your original video.
If you only need the text, copy it into a document you control before removing the recording. If you want to try a different model later, keep the media available: Re-transcribe needs the stored recording.
What should you check if video transcription fails?
Check the audio track, the speech model, and the file format. Video transcription converts the audio track to speech text. A silent video or a file with only music does not contain the spoken information this workflow needs.
| Problem | What to check | Action |
|---|---|---|
| Transcribe file or Voice is unavailable | Pro installation and license access | Complete Pro setup and open the Voice feature. Core chat microphone input is a different workflow. |
| No transcription model is installed | Voice settings > Transcription | Download a model and confirm that it is Active. |
| No speech is detected | The video’s audio track | Play the original and check that speech is audible. Try a shorter sample. |
| File type is unsupported | The actual container and extension | Use a supported video format. Renaming an extension does not convert a file. |
| Drag-and-drop says the file is too large | The displayed 1 GB import limit | Use a shorter clip or an audio-only copy of the section you need. |
| A transcript contains errors | Recording quality and selected model | Compare another model, then use Re-transcribe on the same card. Check the result against the recording. |
| Search shows no matches | Actual transcript spelling and recording age | Search a visible word or part of the title. Check retention settings and the newest-500 display limit. |
| The app reports that media is missing | The stored recording file | Import the original again if you still have it. |
Re-transcribe updates the existing recording’s transcript. Copy the old text first if you want to compare both versions.
Does the video go to a cloud transcription service?
This Voice file-import workflow sends the saved file to a local speech engine. It stages a local copy, decodes the audio, and stores the transcript in the app’s local recording library. You do not need a cloud speech API key for that processing.
Pro activation and app or model downloads are separate setup steps that use a connection. After setup, you can check local processing by disconnecting Wi-Fi and Ethernet before importing a short video, while Pro access is available.
Copying the transcript into a cloud document or sharing the video elsewhere is a separate action. Local transcription does not change where those other applications store data.
Saved-video transcription FAQ
Does this also read text shown on the screen?
This workflow transcribes the audio track. Slides, captions burned into the image, and silent on-screen text need visual analysis. They are not part of the speech transcript.
Can I import the video into a free project instead?
The core project video-import workflow analyzes sampled frames. It does not use the same saved-video speech transcription path described here. Use Pro Voice > Transcribe file when you need the spoken words in the searchable Voice library.
Does the transcript identify every speaker?
This guide does not provide speaker labels. The Voice import produces speech text for the recording. Do not treat it as a speaker-attributed meeting record.
Can I search without a chat model loaded?
Yes. Voice search filters the saved titles and transcripts. It does not ask a chat model to generate an answer.
Start with one short video. Transcribe it, search for a word you can see in the result, and copy the text after you have checked it against the recording.