Ollama, LM Studio, llama.cpp and the others send nothing: the model runs here, and the electricity consumed is this socket's. It is the only case where the grid chosen in the settings applies.
Billed on inference time, 0.75 Wh per minute — about 45 W of machine under load. Bytes do not count: they travel over the loopback interface. Detection works on the process name and CPU load, never on the network.
| Service | Process |
|---|---|
| Ollama (local) | ollama |
| LM Studio (local) | lm studio · lmstudio |
| llama.cpp (local) | llama-server · llama-cli |
| Jan (local) | jan-nitro · nitro |
| vLLM (local) | vllm |
| LocalAI (local) | local-ai |
| GPT4All (local) | gpt4all |
| Msty (local) | msty |
| AnythingLLM (local) | anythingllm |
| ComfyUI (local) | comfyui |
| KoboldCpp (local) | koboldcpp |
| RamaLama (local) | ramalama |