You want an AI model that helps with your work without making the rest of your computer struggle.
OGAD (Off Grid AI Desktop) shows model choices alongside memory guidance. Start with a model that fits your machine, try it on a real task, and move up only when the answer needs it. Local chat is free and works without internet after the required model files are downloaded.
Download OGAD for Mac or Windows

What would you like to do with Off Grid AI?
Have a feature or use case you would like us to support? Tell us what you want to do and which device you use.
Write to support@offgridmobileai.co, join our Slack community, or talk to us on Reddit.
Choose for the task, then for the memory available
A short email draft, a document comparison and an image analysis can need different models. A larger model is not automatically the right first download.
In Models, choose the relevant model type and inspect its details. Where the catalog supplies a Min RAM value, treat that as guidance for the model, not spare memory reserved for your other apps. OGAD can show Tight on RAM or Won’t fit — Load anyway when a model exceeds its fit guidance.
For your first model, choose an option that fits comfortably. The override is useful for an informed experiment; it does not add physical memory.
Download size is not working memory
A 3 GB model download does not mean the entire task will use exactly 3 GB of RAM. The running engine needs memory for the model, the conversation context and other work. Your browser, editor and operating system also need room.
| What you see | What it helps you decide |
|---|---|
| Model download size | Disk space and download cost |
| Min RAM and fit warning | Whether the model is a sensible starting point |
| Context window | How much conversation the model can use at once |
| Answer on your own sample task | Whether that choice is useful for your work |
Quantized models store weights more compactly. They can make local AI practical on smaller machines, but smaller files still need a supported model architecture and enough working memory.
Get a useful first answer
- Open Models in OGAD.
- Pick a text model with memory guidance suited to your computer.
- Download it, then load it.
- Open Chat and give it a small real task, such as rewriting one paragraph or comparing two short notes.
- Check both the answer and whether your usual apps remain responsive.
Use the same task when you compare a second model. A useful model is one that gives adequate answers within the resources you can spare.
When a model is close to the limit
Close apps you do not need. Reduce the Context window in the model settings if a long conversation is consuming too much memory. The app can also reduce effective context to fit its memory budget, so a model’s advertised maximum is not a promise for every machine.
If the model still cannot load, choose a smaller model or smaller supported quantization. Do not keep retrying the same oversized choice and expect a different memory limit.
For private offline work, select a local model. Choosing a remote provider changes where the request is processed.
The controls described here are available in OGAD 0.0.51.
How do you choose a compatible GGUF file?
GGUF is a file format, not a promise that every model in that format will run in the installed engine. Start with a supported catalog model. Open its details and inspect Available GGUF files when that section is present. Compare the quantization and download size, and check whether the chosen file includes a matching vision projector.
Keep three checks separate: a supported model architecture, a complete set of required files, and enough working memory. A filename that looks smaller only addresses part of that decision. For an image-reading model, the matching projector can be necessary even when the text model already runs.
Use the model’s linked repository to check its intended use and license before downloading a different variant. When you are unsure, use the app’s recommended catalog choice for the first test. An arbitrary GGUF file or an unrelated projector is not a compatibility fix.
Compare two choices on a task you actually do
Suppose your normal use is rewriting short client emails. Give the first model a draft containing a date, an unresolved question and a promise you must not change. Ask for a clearer version under 120 words. Check whether those three facts survive.
Then run the same prompt with a second model that also fits the computer. Keep other work similar and separate the first load from a later reply. Choose based on the amount of correction needed and whether you can keep working while it responds.
For document questions, use a different sample. Include the passage that contains the answer and ask for a short quote that supports it. A model that rewrites emails well is not automatically the best choice for careful source-based answers.
Keep a small comparison note:
| Check | What to record |
|---|---|
| Can it load? | Whether it starts without closing essential apps |
| Is the answer useful? | Missing facts, invented details or formatting failures |
| Is it comfortable to use? | Whether ordinary work remains responsive |
| Does the task need more context? | Which source material must fit in the request |
This gives you a reason to download a larger model, if you need one. If both choices handle the task adequately, the smaller one leaves more room for other work. If neither preserves the facts, reconsider the prompt and model capability before spending more disk space on variants.
Start with the model you can use every day
Download OGAD and try one model that fits comfortably. Bring one real task, check the result, and keep the computer usable for the rest of your work.