How to Test Faster Local AI Replies With Speculative Decoding in Off Grid AI in 2026 | Off Grid AI
Did this land?

You have a local model that writes useful answers, but waiting for each long reply breaks your flow. Before downloading another model, you can test whether the one you already use can generate its answer faster.

OGAD (Off Grid AI Desktop) exposes speculative decoding in its text settings. Start with N-gram, which requires no second model. Compare it with Off on a real writing or coding task, then keep the setting only if it reduces the wait for a result you can use.

Get OGAD for Mac or Windows

OGAD desktop chat interface


What would you like to do with Off Grid AI?

Have a feature or use case you would like us to support? Tell us what you want to do and which device you use.

Write to support@offgridmobileai.co, join our Slack community, or talk to us on Reddit.

The control is available without Pro in OGAD 0.0.51. With the app and a working local model downloaded, this comparison can run without internet. Your prompts stay on your computer when you select the local model and do not use remote tools.

What can speculative decoding help you finish faster?

Think about the replies you wait for repeatedly: a revised project update, a rewritten function, or a report built from structured notes. You already know what a useful result looks like. That makes these tasks a better starting point than an unfamiliar benchmark prompt.

The goal is to reduce the time between asking and getting usable output. Speculative decoding changes how tokens are generated. It does not increase your model’s knowledge, add memory, or make an unsupported model fit your computer.

It can be worth testing when the answer has started but the stream of text feels slow. If the main problem is a long pause before the first word, inspect loading and input processing too. A decoding change may leave most of that delay untouched.

How does it work?

A language model normally generates output in successive token steps. A speculative method proposes several possible next tokens, and the main model checks those proposals. Accepted proposals can let generation advance with less repeated work. Poor proposals add checking work without enough benefit.

There are different ways to make the proposal. A separate draft model uses another, smaller model. N-gram methods use patterns in token sequences. OGAD uses the n-gram cache method, which proposes tokens from sequence statistics. They do not require a second model download. The llama.cpp speculative decoding documentation explains these approaches.

This makes a task that reuses existing wording a useful experiment. It does not establish that your particular email, code file, or computer will be faster. Measure the job you actually do.

What do you need before you start?

Use a downloaded local text model that already answers reliably in OGAD. Keep enough free memory for that model and its context. If it frequently fails to load, solve that first with a smaller model or less context.

Choose a fixed prompt, keep the input in a separate note, and record your current settings. Leave Temperature, Context window, and Max output unchanged during the first comparison. Changing the model and the decoding mode together makes the result difficult to interpret.

Close or finish other heavy work that would interfere with the comparison. There is no need to download a second model for the N-gram route below.

Use one complete, repeatable task

Here is a small example you can paste into a fresh chat. The names and dates are sample data, not a reported customer result.

Rewrite the following project update for a client.

Use these headings: Completed, Waiting on, Next steps.
Keep every name and date. Do not add commitments or imply that
unfinished work is complete. Use 120–180 words.

Notes:
- The Atlas onboarding draft was reviewed on 12 September.
- Maya supplied the new screenshots on 13 September.
- The pricing page still needs approval from Leon.
- Leon has not confirmed when approval will arrive.
- The accessibility review found two keyboard-navigation issues.
- Sam will check fixes for those issues on 16 September.
- The next client review is on 18 September.
- We need the pricing decision before preparing the final review pack.

Before timing anything, decide what must survive the rewrite: Leon’s approval is still missing, no approval date is known, Sam’s check is planned, and the client review is on 18 September. A reply that changes those facts needs correction even if it arrives quickly.

After this first check, use your own longer recurring task. A short example helps confirm the controls; it may be too small to reveal a useful speed difference.

Compare Off and N-gram in OGAD

  1. Load the local text model you want to keep using.
  2. Open chat settings and select Text. Turn on Generation details.
  3. Under Advanced (reloads the model), leave Speculative decoding set to Off.
  4. Run a short warm-up request after the model loads. Keep that startup result separate from your comparison.
  5. In a fresh chat, run the fixed prompt. Let it finish, then record the available timing details and check the facts.
  6. Repeat the same prompt in another fresh chat. Keep the surrounding settings unchanged.
  7. Change Speculative decoding to N-gram and allow the model to reload.
  8. Warm up again, then repeat the same fresh-chat comparisons.

Do not send one prompt into a long conversation and the other into an empty chat. The model would receive different amounts of input. Likewise, comparing a first request after loading with a later warm request mixes startup time with the setting you intended to test.

Two or three runs per setting give you more useful evidence than one unusually fast reply. If the results vary widely, repeat under steadier conditions before keeping the change.

Which numbers should you compare?

Open the reply’s generation details. Use the values OGAD actually reports; leave missing values blank. This table is a record for your own results, not a benchmark claim.

Mode and run TTFT (first token) tok/s Output tokens Total time Facts and format correct?
Off, run 1 — — — — —
Off, run 2 — — — — —
N-gram, run 1 — — — — —
N-gram, run 2 — — — — —

Time to first token describes the initial wait. Tokens per second describes output generation. Total time includes the whole request, while output tokens helps you see whether one reply was simply longer.

If a reply finishes sooner but contains much less text, that alone does not show a decoding improvement. If the generation rate improves but total time barely changes, input processing or another part of the request may dominate. For more detail, see how to read OGAD’s reply timing.

When should you try MTP or DFlash?

Start with the option that your current model supports. The menu does not make every method available for every model.

Option What it needs What to check
N-gram No separate model Compare the actual task with Off
MTP Prediction support in the main model The app must recognize that support
Draft model An explicitly compatible draft pair Do not choose an arbitrary small GGUF
DFlash A compatible installed draft file associated with the selected model Complete the matching model-specific download first

In the checked release, ordinary Draft model pairs are not declared in the catalog, so that choice can remain disabled. MTP is disabled when the app identifies the main model as unsupported. DFlash requires the matching installed file; a random draft with a similar name is not enough.

The released OGAD model checks enforce those distinctions. Extra model files also use storage and can need more working memory. Try another method only when it is available for your selected model and you have a reason to compare it.

What should you do with the result?

Keep N-gram if repeated runs finish useful work sooner and the app remains stable. If the difference is small or inconsistent, leaving it off is a reasonable result. There is no benefit in maintaining a tuning choice you cannot distinguish from normal variation.

Result Next action
Generation improves and the output meets the brief Keep the setting and try a second real task
Faster token rate, little change in total wait Check input length and time to first token
No clear benefit Return to Off or compare a more representative task
Loading fails or performance becomes worse Return to Off and confirm the previous setup works
Both modes feel too slow Compare a smaller model or a more focused request

A setting that helps a repeated rewrite may not help an open-ended conversation. Keep the comparison tied to your actual work. If the model itself is too large, use the broader local AI speed guide before adding more complexity.

Try OGAD with one task you repeat every week. Compare Off and N-gram, inspect the answer, and keep the configuration that gets you to a finished draft with less waiting.