You can transcribe a saved audio recording on your Mac without sending it to a cloud service. OGAD (Off Grid AI Desktop) uses a local speech model to turn an attached audio file into text. Download the model first, import the recording, and open its transcript before sending anything to a chat model.
| Download OGAD | Desktop releases |

What would you like to do with Off Grid AI?
Have a feature or use case you would like us to support? Tell us what you want to do and which device you use.
Write to support@offgridmobileai.co, join our Slack community, or talk to us on Reddit.
A long interview or lecture has a different problem from a short voice note. You need the source audio on disk, enough time for transcription, and a way to check that the text reaches the end of the recording.
This guide uses the free desktop app on an Apple Silicon Mac. It covers existing audio files. The separate Pro meeting recorder captures new meetings and has its own controls.
What do you need before you start?
Use an Apple Silicon Mac, M1 or later, with macOS 13 or later. Install OGAD and download a transcription model. Keep the recording in a local folder, with enough free storage for processing. Internet is needed for setup; the local transcription step can run with networking off.
These steps use controls available in desktop stable version 0.0.51. The desktop download page lists supported Macs and available builds.
Before importing a long recording:
- Play the source file to confirm that it opens and contains speech.
- Check that its beginning and ending are present.
- Keep a copy of the original recording.
- Complete the speech-model download while connected.
If the recording is stored in a cloud drive, download it to your Mac first. A file placeholder that needs a network download is not a local offline source.
Which audio formats can you transcribe?
The desktop audio importer accepts MP3, WAV, M4A, AAC, OGG, OGA, Opus, FLAC, AIFF, and AIF files. It decodes the audio locally before passing it to the speech model. A supported extension does not guarantee that an incomplete or damaged recording can be read.
| Your source | What to import |
|---|---|
| An audio interview or podcast recording | Its MP3, M4A, WAV, or other supported audio file |
| A voice-recorder export | A supported audio file saved locally |
| A video meeting recording | Export the audio to a supported audio format first |
| A very long recording that repeatedly fails | Shorter audio parts, saved in order with a local editor |
Do not rename an MP4 file to M4A. Renaming does not convert the contents. The app handles video attachments differently, so use a real audio export for this procedure.
Which speech model should you start with?
Choose a small model for the first import, then compare a larger model if the transcript misses too many words. Whisper Base is an approximately 148 MB download in the desktop catalog. Whisper Small, at approximately 488 MB, is another option once you have checked the workflow.
| Desktop model | Approximate download | Reason to choose it |
|---|---|---|
| Whisper Base | 148 MB | A first pass on a recording or a short sample |
| Whisper Small | 488 MB | A comparison if Base misses important words |
| Whisper Medium | 1.53 GB | Further testing when more memory is available |
| Whisper Large v3 Turbo | 1.62 GB | Another model to compare on the same source |
These sizes describe model files. They do not include the memory required while the model runs or the storage used to process your recording. A larger model can take more resources without resolving problems caused by poor audio.
The local engine is whisper.cpp. The models listed above are multilingual. English-only variants have .en in their filenames.
How do you import the recording and read its transcript?
Select a local transcription model and the spoken language. In Chat, use Attach files to select the saved audio. Wait for the attachment to finish processing, then click its text preview. The full returned transcript opens in a side panel. You do not need to send a message to read it.
1. Download and select the speech model
Open Models > Transcription. Download Whisper Base or another suitable local model. When the download finishes, select it and choose Use this model.
Open Settings for the active model. Under Transcription, check Current model. Make sure it identifies the local model you chose.
2. Set the spoken language
Set Spoken language to the language used in the recording. Use Auto-detect with a multilingual model when needed.
This setting controls speech recognition. The language selected for spoken AI replies is a different setting.
3. Attach one audio file
Open Chat. Select the + control beside the composer, then Attach files. Choose your audio file from the local folder. You can also drop the file onto the composer.
The attachment card shows Processing… while the file is handled. Start with one file so that you can match the result to its source.
4. Open the transcript
When text appears on the attachment card, click its preview. The Click to expand control opens a side panel with the file’s name and the returned text.
Read the transcript there. Select and copy text into a local document if you want to save an edited version. Keep the original audio so that you can check unclear passages later.
This attachment view shows recognized text. There is no need to ask a chat model to recreate the transcript from a summary.
5. Check the beginning, middle, and end
Compare several passages with the source audio. Include the final spoken sentence so that you can detect an incomplete result.
Pay particular attention to numbers, names, and words that change the meaning. An omitted “not” can matter more than several punctuation errors.
How long a recording can you transcribe?
There is no useful unlimited-length promise for this workflow. The app processes each attached audio file as a transcription job. A recording that completes on one Mac or with one model can fail with a larger model, less available memory, or damaged audio.
The local Whisper path has processing time limits. Audio conversion can time out after ten minutes, and a transcription attempt can time out after thirty minutes. Those are time spent processing, not maximum recording durations. A thirty-minute recording is not automatically too long, and a shorter one is not guaranteed to finish.
The attachment importer does not automatically split a long recording into a queue of smaller transcription jobs. If a file repeatedly fails or leaves your Mac short of memory:
- Keep the original file unchanged.
- Use a local audio editor to make shorter parts at pauses in the speech.
- Name the parts in order, such as
interview-01.m4aandinterview-02.m4a. - Import each part separately and save its checked transcript under the same name.
Choose part lengths based on what your Mac completes reliably. If you include a small overlap at each cut, compare the boundary text and remove duplicated words when combining the transcripts.
Can you get timestamps and speaker names from this view?
This audio-attachment workflow returns plain text. Do not expect a timed subtitle file or reliable named-speaker labels from the attachment preview. If the recording contains an interview, check speaker changes against the audio and add names manually when you know who spoke.
The separate meeting recorder has different recording and transcript features. Its export controls do not establish that an ordinary audio attachment has the same output options.
Can you summarize the transcript offline afterward?
Yes, with a downloaded local text model. Transcription and summarization are separate steps. The speech model returns the words; the text model produces an answer about them. Keep both model choices local if you want the complete workflow to work without internet.
After checking the transcript, you can ask for a summary based on the attachment. For a long interview, try:
Summarize the main points in this transcript. Keep decisions separate from suggestions. If a name, date, or amount is unclear, say so.
A long transcript can exceed the text model’s available context. If the response omits sections, work through smaller parts. Keep the source transcript alongside each summary so that you can check the answer.
What should you do when an import fails?
| What happens | What to check or change |
|---|---|
| No transcription model is found | Finish downloading a local speech model and make it active. |
| The card reports a file-reading error | Confirm that the source plays, has an audio track, and uses a supported audio format. |
| A video produces descriptions of frames | Export its audio with a local tool, then attach the audio file. |
| A long file fails repeatedly | Try a smaller speech model or divide the source into shorter audio parts. |
| The language is wrong | Set Spoken language explicitly and use a multilingual model. |
| The output is empty | Check that the file contains audible speech and that the correct recording was selected. |
| The transcript works but a summary does not | Check the selected text model and its available context separately. |
Avoid repeatedly importing the full recording before you know why it failed. A short section from the same file can help you distinguish a format problem from a long-job problem.
How do you check that no cloud upload is needed?
Complete the setup and keep the source audio on disk. Then turn off Wi-Fi and disconnect Ethernet or any other network connection. Import a short audio file with the local transcription model selected. If the text appears, repeat with your longer file or its first part.
The desktop file picker passes the recording to the app on your Mac. With a local speech model active, recognition happens there. The app also supports remote transcription services, so check the selected model before importing private audio.
Sending a transcript to a remote chat model, sharing it, or saving it in a synced folder is a separate action. Local transcription does not decide what you do with the text afterward.
Start with one supported audio file. Open its transcript before sending anything, check the final sentence, and save the text you have reviewed.