AI News · New apps ·

CIQ releases Fuzzball 4.3 for running open-weight models on customer GPUs

CIQ releases Fuzzball 4.3 for running open-weight models on customer GPUs

CIQ, the enterprise software company behind Rocky Linux, announced general availability of Fuzzball 4.3. The release adds ready-to-run AI models and agents to its sovereign AI orchestration platform. Customers can deploy validated open-weight models from a catalog on their own GPUs with one command. CIQ positions the release around production control and governance on existing infrastructure.

Key points

  • CIQ announced general availability of Fuzzball 4.3 for deploying models and agents on customer infrastructure.
  • The catalog includes 11 presets, with one-command deployment and support for other compatible Hugging Face models.
  • Models can scale to zero when idle and run across multiple nodes.
  • Fuzzball supports NVIDIA and AMD GPUs across on-premises infrastructure and several cloud platforms.
  • Pricing, performance benchmarks and measured cost savings were not reported in the announcement.

What happened: CIQ, the enterprise software company behind Rocky Linux, announced general availability of Fuzzball 4.3 on October 8, 2026. The release adds ready-to-run AI models and agents to its sovereign AI orchestration platform, which manages workloads on infrastructure customers control. In a company announcement distributed by GlobeNewswire, CIQ said customers can launch validated open-weight models from a catalog on their own GPUs with one command. The pitch is to simplify production deployment without requiring organizations to redesign their existing environments.

The details: The self-service catalog contains 11 presets covering models from the GPT-OSS, Llama 4, Gemma 4, Qwen3-Coder, Mistral and Nemotron 3 families. A separate entry based on vLLM, software for serving AI models, supports other compatible Hugging Face models. According to CIQ, models automatically scale behind a single, stable endpoint, the address applications use to access them. They can scale down to zero running instances when idle, releasing GPU capacity for other workloads, and restart for additional requests. Large models can also run across multiple nodes to access more resources.

The details: Fuzzball 4.3 also handles connections between models and agents. CIQ said OpenCode and hermes-agent, two coding agents in its workflow catalog, run inside the cluster and discover active models through gateways without manual integration. Each model includes a LiteLLM gateway with an OpenAI-compatible endpoint. Customers can instead operate one central gateway for all models a user is authorized to access. Another catalog entry offers a retrieval service that connects private agents to enterprise documents.

Who it affects: The release targets teams moving from AI experiments to production on private or existing infrastructure. CIQ frames the operational challenge as assembling infrastructure and manually connecting models, gateways and agents before useful work can begin. Fuzzball supports mixed NVIDIA and AMD GPU fleets, with serving and health-aware scheduling under one operating model. Deployment options include AWS, Google Cloud, Oracle Cloud, CoreWeave and Azure, along with on-premises clusters built using Warewulf, VMware or bare metal. Those options are central to CIQ's emphasis on fitting AI into infrastructure decisions customers have already made.

What to watch: CIQ said administrators can control which organizations and groups use particular resources, idle pools wake only for authenticated requests, and all workloads run unprivileged and rootless, without elevated system permissions. For buyers, the practical question is whether the validated catalog, access controls and deployment options match their models and infrastructure requirements. The packaged approach could reduce manual setup, but performance benchmarks, measured cost savings and pricing were not reported in the announcement. Teams evaluating the release will need those details to assess its operational and financial fit.

Our take

Packaged deployment could reduce the operational work involved in self-hosting AI. Teams should compare the validated catalog and governance capabilities with their model and infrastructure requirements.

Sources