Your Linux computer already has an NVIDIA GPU. You want to put it to work on an AI model, not spend your first session compiling inference engines or choosing CUDA libraries.
OGAD (Off Grid AI Desktop) provides an optional NVIDIA download inside its Linux app. Install the pack, restart, and use a downloaded local model for chat, images or transcription. You can see the backend that actually loaded and use CPU or Vulkan when CUDA cannot run.
This guide covers v0.0.54-beta.108, a prerelease for Ubuntu 24.04 or newer on x64. It uses the core Linux app. Pro is not bundled in this release. Initial app, model and GPU downloads need internet; the local tasks described here can run without it afterward. Release.
| Download OGAD for Linux | Off Grid AI |

What would you like to do with Off Grid AI?
Have a feature or use case you would like us to support? Tell us what you want to do and which device you use.
Write to support@offgridmobileai.co, join our Slack community, or talk to us on Reddit.
What do you need on the Linux computer?
Use the released x64 AppImage or amd64 .deb. There is no Linux ARM package in this release. Stay within the documented Ubuntu version when following this guide; other distributions need their own compatibility checks.
You need a working NVIDIA driver and enough memory for the chosen model. OGAD’s pack supplies runtime components. It does not install the host graphics driver, expand your VRAM, or download every model.
You do not need Docker or a full CUDA developer toolkit to use the installed application. Those requirements in the repository apply to building the Linux application from source. The released pack contains its runtime libraries. Linux installation and build documentation.
The optional Linux pack is about 5.88 GB to download. Leave additional disk space for extraction. The AppImage is a separate download of about 1.08 GB, and model weights are separate again. Starting with one small model keeps the first setup manageable.
Which local AI tasks use the pack?
The pack includes NVIDIA engines for chat and vision, images and speech-to-text. You select processing preferences separately for each kind of task.
| Task | Example | Where to check the preference |
|---|---|---|
| Chat and vision | Turn notes into a useful summary | Model settings → Text |
| Images | Create a draft illustration from a prompt | Model settings → Image |
| Transcription | Turn a spoken question into text | Model settings → Transcription |
The shared download panel also lists Computer Use engine components. That list does not mean the Linux package contains Pro features. This guide stays with the released core workflows.
Spoken replies and search embeddings have separate processing settings. Do not judge the whole application by the backend of one chat request. A working CUDA text model can coexist with CPU speech.
How do you install the NVIDIA components?
1. Get one local model working
Install the Linux beta and complete its setup flow. Settings → Setup & health lets you configure local models and inspect component health. Choose a model suited to the available memory.
Before adding GPU components, make a short local chat request. This gives you a simple baseline: the application can load a model and return an answer. If that does not work, resolve the model download or memory issue first.
2. Download the optional pack
Open Settings → GPU performance. In Use your NVIDIA GPU, choose Download. You can pause and resume the download or continue using the app while it runs.
OGAD checks the archive’s expected size and SHA-256 hash before installation. Wait for it to complete, then use Restart app. The new native engines become available after the restart. Published pack metadata.
If the panel does not detect an NVIDIA driver, check your host driver installation first. The app’s Linux presence check looks for an NVIDIA device; it is not a complete compatibility test. A detected driver does not guarantee that every CUDA engine will load.
3. Load a model with the chosen backend
In Chat, select the Settings button to open Model settings, then choose Text. Use Auto for the normal selection order, or choose CUDA to prefer the NVIDIA route. Backend changes apply when the model next loads, so reload an already-running model before you judge the change.
Read the Now value after loading. The preference is your requested route; Now describes the runtime state. If CUDA cannot run, a non-CPU preference can fall back to another supported route. Explicit CPU selects the CPU path.
What is a useful first GPU task?
Use something small enough to review and repeat. For example, give the local model these notes:
We tested the export on Ubuntu. Two filenames need correction. The documentation draft is ready. The next review is Thursday.
Ask:
Make a short handover with “Done,” “Still needed,” and “Next review.” Use only the facts in these notes.
Check both the result and the active backend. A successful CUDA load is useful only if the chosen model also does the task well. If the answer invents details, improve the prompt or try a different model; changing the GPU backend is not a factual-accuracy setting.
You can then repeat the same request with CPU selected. Reload the model between changes and keep the prompt unchanged. Do not compare the first cold load against a later warm request as if they were the same operation.
For image generation, start with one simple illustration at a modest image size. Image processing has its own backend control. A text model using CUDA does not prove the image engine uses it too.
What if CUDA fails or local AI is still slow?
| Problem | A useful next check |
|---|---|
| CUDA says the download is required | Finish the pack installation and restart the app |
| The preference changed but the active backend did not | Reload the model and inspect Now again |
| A large model fails while a small one works | Check available memory and reduce model size |
| CUDA fails but CPU works | Inspect the request error and the host NVIDIA driver |
| Voice playback uses CPU | Speech has a separate engine and fallback order |
| A Pro workflow is missing | The Linux beta contains core features only |
Open Settings → AI activity for the failed request. The record can show its model, status, backend and error. Change one relevant setting, retry the same small request and compare the new record.
The release supplies runtime components, but it does not establish a universal driver minimum, supported-card list or speed multiplier. CPU and Vulkan are useful working paths when a CUDA route is unavailable. Keep a working configuration while you investigate, rather than making several simultaneous changes.
Keep the first setup focused
A local GPU is valuable when it helps you finish a task on hardware you own. Start with one model, one prompt and a check of the active backend. Add image generation or transcription when you have a reason to use them.
Download the Linux beta, prepare a local model and let the app install the optional NVIDIA components for you.