How to Use AI Running on Your Own Computer From Python in 2026 | Off Grid AI
Did this land?

You can add AI to a Python script without sending the prompt to a cloud provider.

OGAD (Off Grid AI Desktop) runs the model on your computer and exposes an HTTP gateway. Python can call it with the standard library, so the first example needs no AI SDK or provider API key. Download a suitable local model first, then keep inference on the same machine.

Download OGAD for Mac or Windows

Off Grid AI


What would you like to do with Off Grid AI?

Have a feature or use case you would like us to support? Tell us what you want to do and which device you use.

Write to support@offgridmobileai.co, join our Slack community, or talk to us on Reddit.

Prepare one local model

Install OGAD, open Models and download a text model that fits the computer. Select it and try a short prompt in Chat. Then open Gateway and check the local address.

The usual address is http://127.0.0.1:7878. Use the port shown by your running app if it differs. Core local chat and the gateway do not require Pro.

This example turns rough progress notes into a short update. It discovers the selected local chat model instead of guessing the name of a downloaded file.

Send the first request

Save this as local_update.py and run it with Python 3:

import json
from urllib.request import Request, urlopen

BASE = "http://127.0.0.1:7878"

def request_json(path, payload=None):
    data = None if payload is None else json.dumps(payload).encode("utf-8")
    request = Request(
        BASE + path,
        data=data,
        headers={"Content-Type": "application/json"},
    )
    with urlopen(request, timeout=180) as response:
        return json.load(response)

models = request_json("/v1/models")["data"]
model = next(
    (item for item in models
     if item.get("kind") in ("chat", "vision") and not item.get("remote")),
    None,
)
if model is None:
    raise SystemExit("Select a downloaded local text model in OGAD first.")

result = request_json("/v1/chat/completions", {
    "model": model["id"],
    "messages": [{
        "role": "user",
        "content": (
            "Turn these notes into a short update. Keep uncertainty explicit. "
            "Notes: import complete; tests in progress; launch date undecided."
        ),
    }],
    "max_tokens": 160,
    "stream": False,
})
print(result["choices"][0]["message"]["content"])
python3 local_update.py

The expected result is a generated update printed in the terminal. Its exact wording depends on the selected model. Check that it preserves the undecided launch date before reusing it.

Put the model to work on your own data

Replace the sample notes with text your script has already read. Useful first tasks include turning a few changelog notes into a draft, explaining one function or creating a short label for an internal document.

Keep the first input small. Reading a large file into Python does not mean the model can fit the whole file in its context. For a longer document, extract the relevant section or build a retrieval step first.

A system message can set the format you need. Still validate the output before treating it as a date, a command or structured data. The HTTP response shape is predictable; the content generated by the model is not guaranteed to follow every instruction.

Handle ordinary failures

A connection error usually means the app is not running or the port differs. An empty local model selection means you need to select and load a suitable model. A timeout can mean a slow first load, an oversized task or a model that cannot fit; check the app before blindly retrying.

For a production script, catch HTTP errors, log the error body and set a retry policy that fits the task. This small example lets an error stop the script so it cannot quietly substitute an invented result.

The gateway listens on network interfaces and its inference endpoints do not require an API key. These examples use 127.0.0.1 on the same computer. Keep the host on a trusted network and do not expose this port to the public internet.

These API routes are present in OGAD 0.0.51. The running gateway also serves its API reference at /docs.

Turn the sample into a useful file-processing step

After the built-in sample works, read a short UTF-8 note from a file. Add from pathlib import Path, then use notes = Path("notes.txt").read_text(encoding="utf-8") before building the request. Put notes into the prompt where the sample facts are now.

For the first run, use a note you wrote and keep it short enough to inspect. Tell the model what must survive the rewrite: names, dates, unresolved questions and the distinction between completed and pending work. A useful prompt is more specific than “summarize this.”

Check the returned value before saving it. For example:

answer = result["choices"][0]["message"].get("content")
if not isinstance(answer, str) or not answer.strip():
    raise RuntimeError("No final answer text was returned.")
output = Path("update-draft.txt")
with output.open("x", encoding="utf-8") as handle:
    handle.write(answer.strip() + "\n")

This uses exclusive creation, so an existing draft is not silently overwritten. Choose another output name for a second run. The saved file is a draft to review, not a report the script has proved correct.

Decide what success means before adding a loop

For the progress-note example, the launch date must remain undecided and testing must remain in progress. Read the output against those checks. A fluent paragraph that says “ready to launch” has failed the task.

Only after one item works should you process more files. Keep a record of which input produced each draft, handle failures separately and avoid sending many requests at once to a model that already fills the computer’s memory. A clear failed item is easier to recover than an empty file that looks complete.

Add one local AI step to your script

Download OGAD, run the sample and replace the notes with one small piece of your own work. Keep the processing local and the result easy to check.