LivoPC team · AI-assisted editorial summary
Ollama or LM Studio running slowly: how to investigate on your PC
Separate loading, context and text generation, check available acceleration and investigate slow local AI before buying hardware.
Updated on · Games and apps

What to consider first
If Ollama or LM Studio is slow, first identify where the delay occurs: loading the model, processing the prompt or generating the response. Record the model, quantization, context and version; then check whether execution uses the CPU, GPU or both. Repeat a short task with one change at a time to compare equivalent conditions.
This is an inference troubleshooting guide, separate from planning a PC for local AI. GPU usage is not tokens per second; memory usage and response quality are different questions.
For anyone already running a local model who wants to understand a delay or change in behavior.
Watch CPU, GPU, RAM and available sensors during the response with Monitor for Windows. Local use requires no account; sending data to the platform is separate. Coverage depends on hardware, drivers and the installed version.
1. Locate the delay and choose a comparison task
Note the exact model name and variant, quantization, app and its version. Use a familiar short prompt with the same output limit. Compare the first run with another while the model remains loaded. Downloading, loading, prompt evaluation and generation have different costs; combining them obscures the problem you want to solve.
- In the Ollama API, load_duration, prompt_eval_duration and eval_duration help separate the phases; durations are reported in nanoseconds.
- When available, get tokens/s from the inference tool's statistics. Monitor tracks PC resources; it does not measure tokens/s or model quality.
2. Check where the model was loaded
While the task is active, run ollama ps and check PROCESSOR: CPU, GPU or a split between them. This shows model placement, not chip utilization. In LM Studio, check the model's load parameters and GPU offload. Check the exact GPU, operating system and driver against the backend documentation before assuming acceleration is active.
3. Review context and concurrency before replacing parts
A long conversation and simultaneous requests can increase memory requirements. Check the context actually allocated rather than assuming another version's default. To investigate, start a short conversation, stop inference tasks you recognize and repeat. Reducing context changes how much content the model can consider; also check that the response remains suitable for your work.
| Situation | Suggested comparison | What to record |
|---|---|---|
| Only the first response takes a long time | Newly loaded model versus one already loaded | Loading time separately from generation |
| Long conversations become slow | New conversation with the same request | Context and available memory |
| Several requests are waiting | A single isolated request | Queue, concurrency and loaded model |
4. Match readings to the same time interval
Open Monitor and watch available resources as you repeat the request. Record when waiting and generation start, along with other known tasks running at the same time. For dedicated memory and backend details, also use the inference or manufacturer's tool. An unavailable temperature remains unknown; a utilization spike does not confirm a hardware shortfall.
5. Choose the next test based on the result
If the backend does not recognize the GPU, investigate support and drivers. If memory limits the task, compare a smaller variant or another quantization, keeping the request the same and evaluating the response. If the delay only occurs during loading, check how long the model stays in memory. Save versions, settings and results before planning an upgrade; the build guide covers that decision separately.
Frequently asked questions
Does GPU activity mean the entire model is on it?
Not necessarily. Check the placement reported by the backend and available memory. Other tasks can also use the GPU.
Does lower-bit quantization always make it faster?
There is no guarantee. Format, backend and hardware affect the result, and the response may change. Compare the exact variant on your task.
Does Monitor speed up Ollama or LM Studio?
It tracks available resources. It does not change the backend, optimize the model or measure agents, tokens/s or response quality.
How this guide was prepared
AI-assisted editorial summary, with official sources checked on October 11, 2026 and references for each step. Scenarios are educational; no benchmark or training run was performed. The latest evidence on October 11 at 00:14 UTC showed Monitor 1.11.1 undergoing certification; this does not establish that the update is available. No human language review was performed.
- LivoPC Monitor — features and Windows coverage · accessed on
- Ollama — generation statistics · accessed on
- Ollama — loading, memory and concurrency · accessed on
- Ollama — context and allocated memory · accessed on
- Ollama — hardware and driver support · accessed on
- LM Studio — per-model parameters · accessed on
- LM Studio — loading and memory estimation · accessed on
Apply it to your PC
Watch CPU, GPU, RAM and available sensors during the response with Monitor for Windows. Local use requires no account; sending data to the platform is separate. Coverage depends on hardware, drivers and the installed version.
Suggest a guide correction