AI News · Good news ·

llama.cpp adds half-precision support for two WebGPU operations

llama.cpp adds half-precision support for two WebGPU operations

The open-source model runner llama.cpp has added half-precision support in release b11382, according to Oossa. The update introduces f16 support for WebGPU fill and set_rows operations. The reported change concerns those specific operations rather than a new model release.

Key points

  • Oossa reported that llama.cpp release b11382 adds f16 support for WebGPU fill and set_rows operations.
  • The change concerns model-running infrastructure, not a new model release.
  • Benchmarks, measured speed gains and memory savings for this update were not reported.
  • Teams should check whether the two operations affect their deployment before expecting broader improvements.

What happened: The open-source model runner llama.cpp has added half-precision support for two WebGPU operations in release b11382, according to Oossa. The outlet reported the change on October 4, 2026. The update adds f16 support for operations named fill and set_rows. It is a targeted change to the software used to run models, rather than the introduction of a new AI model or a reported improvement in model capabilities.

The details: Oossa describes f16 as half-precision support. The important boundary is that the reported addition applies to fill and set_rows within WebGPU, not to every operation in llama.cpp. For readers evaluating AI software, this distinguishes an infrastructure update from a model launch: the news concerns how the runner supports particular operations. Benchmark results, measured speed improvements and memory savings for b11382 were not reported, so there is no reported basis for quantifying a deployment-wide benefit.

Background: The release follows other llama.cpp updates covered by Oossa, but their benefits should not be attributed to this change. On October 3, the outlet separately reported OpenVINO 2026.4.1 support and fixes to mixture-of-experts model handling. It also reported that release b11372 reduced memory use for the Qwen4exp indexer without a speed loss in long-context runs. Those reports concern different parts of the runner. They do not establish that the new WebGPU half-precision support delivers similar gains.

Who it affects: The immediate audience is teams using llama.cpp for WebGPU-based inference, meaning teams running models through that part of the software. For business teams choosing or maintaining an AI deployment, the relevant question is whether fill and set_rows affect their own setup. The reported change alone does not establish a reason to switch runners or expect broader performance improvements. Its practical importance depends on whether the newly supported operations matter to the deployment being assessed.

What to watch: Before treating b11382 as a performance upgrade, check whether the two operations are relevant to your workload and evaluate the release in that context. A hardware compatibility breakdown and workload-specific results were not reported. For now, the confirmed development is narrow: llama.cpp has expanded half-precision support for two named WebGPU operations, while the effect on complete inference workloads remains unreported.

Our take

This is a targeted infrastructure update for teams using WebGPU-based inference. Check whether these operations affect your deployment before expecting broader performance improvements.

Sources