The process exists, but no single document explains it. Some steps are in notes, some are visible in screenshots, and the exceptions are in a recorded explanation.
OGAD (Off Grid AI Desktop) can help you turn those sources into a process draft with local text, vision, and speech models. Check each source first, then combine the verified details into ordered steps. After setup, the analysis can run on your computer without a cloud AI provider.
| Download OGAD | Desktop releases |

What would you like to do with Off Grid AI?
Have a feature or use case you would like us to support? Tell us what you want to do and which device you use.
Write to support@offgridmobileai.co, join our Slack community, or talk to us on Reddit.
For a small team documenting a repeated task, this makes existing material useful without pretending that a few screenshots capture the entire workflow. The person who knows the process still checks the final instructions.
What can each source tell you?
Use each source for the information it actually contains. Notes can explain intent, screenshots can show visible controls, and audio can preserve spoken instructions. None automatically supplies everything missing from the others.
Suppose you are documenting how a team prepares a client workshop. Your notes explain the purpose, screenshots show the participant setup screen, and the recording explains an exception for external trainers.
| Source | Useful information | Common gap |
|---|---|---|
| Notes | Decisions and reasons | Exact interface steps |
| Screenshots | Visible labels and state | What happened before or after |
| Recording | Spoken explanation and exceptions | Silent clicks and visual details |
| Final work product | Expected result | The route used to create it |
Keep a source label with every important instruction so you can return to the evidence during review.
What do you need to set up?
Install OGAD and download the local models needed for your sources. Text drafting uses a text model, screenshots need a compatible vision model, and audio transcription needs a speech model. One downloaded model does not necessarily cover all three jobs.
The workflow uses free core features on supported Mac and Windows computers. 0.0.54-beta.108 also includes Linux beta packages. Complete app, model, and runtime downloads while connected.
The speech-setting steps below use Mac; setup can differ on Windows. Start with a short process and a few clear source files so you can check the result before handling a larger collection.
How do you prepare the written and spoken material?
Use readable notes such as TXT, Markdown, DOCX, or text PDF. For audio, select a local transcription model in Models, then choose Current model and Spoken language under Transcription in model settings.
Open chat and use + > Attach files to add the audio. Wait for processing and inspect the text preview. Copy the transcript into a local editor for corrections, keeping unclear terms marked until you check the recording.
If you have video, export audio for this step. The ordinary video attachment path reads sampled frames rather than transcribing the complete soundtrack.
Do not infer a missing instruction because it seems like the usual way to perform the task. Mark it as a question for the process owner.
How do you use the screenshots?
Select a compatible local vision model and add one image through + > Add image. Ask for visible labels and state, not an imagined sequence of clicks.
Describe the visible controls and relevant state in this screenshot. Copy readable labels exactly and mark uncertain text. Do not infer actions that happened outside the image or claim that a hidden menu contains a particular option.
The chat image path sends image files to the multimodal model. It does not turn a screenshot into access to the live application.
Check the output against the picture. If a label is small or blurred, inspect it yourself or take a clearer screenshot. Save the checked observation with the source image name.
How do you combine the sources into steps?
Give the local text model the checked notes, transcript, and screenshot observations. Ask it to distinguish supported steps from missing information before drafting a finished procedure.
Build a process outline from these checked sources. For each step, state the action, required input, expected result, and supporting source. Keep exceptions with the relevant step. Mark gaps or conflicting instructions as “Needs confirmation.” Do not invent missing controls or permissions.
For the workshop example, the external-trainer exception should appear where access is configured. It should not disappear into a general notes section the reader might miss.
Review the proposed order with someone who knows the task. A screenshot may show the end of a step, while the recording explains an earlier prerequisite. The model can misorder those details even when each one is individually accurate.
What makes the procedure useful for a new reader?
Name what the person needs before starting and how to tell the work is complete. Keep each step focused on an action the reader can perform. Explain the condition for any branch or exception.
A useful structure is:
- Purpose and expected result.
- Required access, files, or information.
- Ordered actions.
- Exceptions and stop points.
- Completion checks.
- Source or owner for unresolved questions.
Ask the model to flag vague wording such as “use the usual settings” or “check everything.” Replace it with verified details, or keep it as a question until the owner supplies the answer.
How do you validate the draft?
Have a suitable colleague follow the procedure in an appropriate test context. Record where they need an explanation or encounter a different screen. Update the instructions from what you observe.
The draft should not be treated as validated merely because it reads clearly. A step can sound precise while using a label from an old screenshot.
| Issue | What to check |
|---|---|
| A control cannot be found | Source version and current interface |
| A required input is missing | Prerequisite notes and process-owner confirmation |
| Two sources disagree | Dates, conditions, and approval status |
| The final result is unclear | The actual completed work product |
| A safety-critical step is inferred | Stop and obtain the relevant authoritative procedure |
Keep a version date and a named review owner in your own document process. The app does not automatically detect every later change in the system being documented.
Document one repeated task
Download OGAD and gather a small set of sources for one process. Check each source, combine the verified details, and test the draft before sharing it.
Keep text, vision, and speech selections local for this workflow after setup. Later publication, file sharing, and remote models have separate connections. The useful result is a procedure grounded in the work your team actually does.