News
CAPACITY TEST:

Hugging Face Adds 4-Bit Nunchaku Loading To Diffusers

Newsroom brief

Hugging Face added Nunchaku Lite support to Diffusers, letting developers load 4-bit diffusion checkpoints with `from_pretrained()` while using Hub-delivered CUDA kernels.

Verified against source materialEdited by SendTech Times Chips & Compute Desk
Hugging Face Adds 4-Bit Nunchaku Loading To Diffusers
Image source: Hugging Face

Diffusers gets a 4-bit loading path

A July 23 Hugging Face blog post introduced a Diffusers integration for Nunchaku Lite, giving text-to-image developers a way to load Nunchaku-style 4-bit checkpoints through the standard `from_pretrained()` workflow.

The post describes modern BF16 text-to-image pipelines as often requiring 20-30 GB of VRAM, while earlier Diffusers quantization options such as bitsandbytes, GGUF, torchao and Quanto mainly reduce model weight storage.

According to the same technical note, SVDQuant runs the main diffusion-transformer layers with 4-bit weights and activations, or W4A4, so the denoising loop can use less memory and run faster.

Kernels arrive through the Hub

Nunchaku Lite loads without a custom pipeline class, separate inference engine or local CUDA compilation, according to the post.

The runtime patches relevant `nn.Linear` modules in a stock Diffusers model with SVDQ or AWQ linear layers before loading the checkpoint, and CUDA kernels are downloaded through the `kernels` package.

The integration uses `svdq_w4a4` layers for attention and MLP projections, with INT4 and NVFP4 variants.

A second `awq_w4a16` path covers precision-sensitive normalization and modulation components in model families including FLUX and Qwen-Image.

The hardware table says NVFP4 checkpoints require NVIDIA Blackwell hardware, including RTX 50 series, RTX PRO 6000 and B200 GPUs, while INT4 variants support Turing, Ampere and Ada devices such as RTX 30 and 40 series cards, A100 and L40S.

Benchmarks stay vendor-owned

According to Hugging Face's example, one Nunchaku NVFP4 transformer paired with a bitsandbytes NF4 text encoder generated a square 1024-pixel image in roughly 1.7 seconds on an RTX 5090, while peak memory was about 12 GB; the same paragraph placed the BF16 pipeline at about 24 GB.

In a separate RTX PRO 6000 Blackwell benchmark, the blog listed a BF16 baseline at 3.00 seconds for the full pipeline and 31.1 GB peak VRAM.

The blog's benchmark table listed Nunchaku Lite NVFP4 at 2.27 seconds and 20.6 GB peak VRAM, Nunchaku Lite with `torch.compile` at 1.68 seconds and 20.6 GB, and Nunchaku Lite with an NF4 text encoder at 2.29 seconds and 16.0 GB.

The post characterized the result as up to 50% lower peak VRAM, roughly 30% better latency and as much as a 1.8x speedup with `torch.compile`; those figures remain vendor benchmark claims rather than independent lab measurements.

Quantization workflow widens

The post also points developers to `diffuse-compressor`, a companion toolkit for calibrating, quantizing, packaging and publishing new Diffusers models.

The FLUX.2 Klein 4B example workflow says the inspection stage should identify 100 SVDQ targets, 3 AWQ targets and 6 dense outer linears before quantization proceeds.

The remaining boundary is hardware and model structure.

The compatibility note excludes Volta and Hopper GPUs from current 4-bit kernel support, and the generic Nunchaku Lite path does not infer architecture-specific fused rewrites such as combined QKV projections.

That makes the integration useful as a standard Diffusers loading route, while specialized engines may still matter when a model family needs fused modules beyond the generic layer replacement path.

The post did not disclose production customer names for the new Diffusers path, so deployment proof is still limited to the published benchmark examples.

Share this article
inXf

Related articles

More
Micron And Anthropic Put AI Memory Supply Inside Claude’s Compute Plan
Chips & Semiconductors

Micron And Anthropic Put AI Memory Supply Inside Claude’s Compute Plan

Micron and Anthropic signed a strategic agreement covering AI memory and storage architecture, supply, Claude adoption inside Micron and a Series H investment.

AMD Linux Patch Adds Low-Power CPU Core Support For Future Chips
Chips & Semiconductors

AMD Linux Patch Adds Low-Power CPU Core Support For Future Chips

AMD has submitted Linux kernel patches that add a low-power CPU core classification beside performance and efficiency cores, The public record still lacks the first processor, launch date or product line that will use the new core type.

Samsung Sets HBM And Server SSD Efficiency Targets With Fab Water Gap Still Open
Chips & Semiconductors

Samsung Sets HBM And Server SSD Efficiency Targets With Fab Water Gap Still Open

Samsung’s 2026 Sustainability Report says lower-power HBM and server SSD targets support AI infrastructure, while its semiconductor division still has a 2050 net-zero timetable and limited product-level commercial detail.

Intel Dunlow Manifests Point To 28-Core Nova Lake-S Xeon Platform
Chips & Semiconductors

Intel Dunlow Manifests Point To 28-Core Nova Lake-S Xeon Platform

Tom's Hardware reported that NBD shipment manifests point to an Intel Dunlow workstation and entry-server platform with up to 28 cores, LGA1954 packaging, dual-channel memory and 95W processor base power. The details remain pre-launch documentation, and launch dates, pricing, core mix and customers remain outside Intel confirmation.

Samsung Starts PM1763 PCIe 6.0 SSD Production For AI Servers
Chips & Semiconductors

Samsung Starts PM1763 PCIe 6.0 SSD Production For AI Servers

Samsung has started mass production of the PM1763, a PCIe 6.0 enterprise SSD for AI and HPC servers with 28,400 MB/s sequential reads in a 16TB configuration, according to ServeTheHome. Customers, pricing, shipment volumes and server qualification dates remain outside the report.

Infineon Opens €5 Billion Dresden Fab Three Months Early Without Customer Names
Chips & Semiconductors

Infineon Opens €5 Billion Dresden Fab Three Months Early Without Customer Names

Infineon Technologies opened its €5 billion Module 4 smart power fab in Dresden three months ahead of schedule. Customers, order volumes or utilisation targets remain outside the public record.

Memory Prices Push US PC Shipments Down 7%
Chips & Semiconductors

Memory Prices Push US PC Shipments Down 7%

Omdia data cited by Tom Hardware showed US PC shipments fell to 15.8 million units in the first quarter of 2026, as memory and storage chip shortages hit entry-level laptops and pushed the market toward a forecast 14.4% contraction.

Arm and Supermicro Put Agentic AI Servers to a CPU Test
Chips & Semiconductors

Arm and Supermicro Put Agentic AI Servers to a CPU Test

Supermicro has introduced new server platforms built around Arm’s AGI CPU for inference-heavy and agentic AI workloads across cloud, enterprise and edge deployments. Arm says the AGI CPU includes up to 136 Arm Neoverse V3 cores, 12 DDR5 memory channels running at up to 8800 MT/s and PCIe Gen6 connectivity within a 300W power envelope. The key test is whether operators can use these CPU-heavy designs to add inference capacity without creating new pressure on power and cooling.

Keep Reading

More Stories

Latest
AI Distillation Debate Moves From Labs To WashingtonAIJul 27, 2026AI Distillation Debate Moves From Labs To WashingtonCNBC reported that AI distillation has become a policy fight after Moonshot AI's Kimi K3 raised questions about open-weight models, proprietary model output and U.S. restrictions.CFTC Warning Narrows Prediction-Market Contract FilingsFintech & Digital PaymentsJul 27, 2026CFTC Warning Narrows Prediction-Market Contract FilingsCoinDesk reported on July 26 that the CFTC warned prediction-market operators against broad template certifications, tightening the filing burden for event contracts while courts still test the regulator’s authority.B Capital Names Ex-G42 AI Chief To Doha Investment RoleAIJul 27, 2026B Capital Names Ex-G42 AI Chief To Doha Investment RoleB Capital has appointed former G42 executive Dr Andrew Jackson as General Partner and Chief AI Officer in Doha, tying the firm’s Gulf office to AI investment governance and Stargate UAE experience.Ajman AI Agent Renews Trade Licences Through AjmanOneEconomyJul 27, 2026Ajman AI Agent Renews Trade Licences Through AjmanOneMiddle East AI News reported that Ajman used agentic AI to renew a trade licence through AjmanOne, moving a government service from chatbot support toward an automated transaction flow that still depends on governed data and system integration.AMD Helios AI Racks Bring Epyc CPUs To Nvidia Compute ChallengeChips & SemiconductorsJul 26, 2026AMD Helios AI Racks Bring Epyc CPUs To Nvidia Compute ChallengeData Center Knowledge reported that AMD moved Helios into production with MI455X GPUs, Epyc processors, Pensando networking and ROCm software, while vendor performance claims still lack third-party benchmark validation.Dubai Airports Adds Smart Gate Pre-Check Before Summer PeakEconomyJul 26, 2026Dubai Airports Adds Smart Gate Pre-Check Before Summer PeakEconomy Middle East reported that Dubai Airports launched a Smart Gates Eligibility Pre-Check at DXB, letting passengers confirm eligibility through Pocket Flights or terminal QR codes before passport control.Hong Kong Indoor 5G Gap Shows Building-Level Network DivideTelco & ConnectivityJul 26, 2026Hong Kong Indoor 5G Gap Shows Building-Level Network DivideA Speedtest Intelligence study says Hong Kong’s indoor 5G experience varies sharply by building, putting the focus on shared in-building systems rather than outdoor coverage alone.CIRCIA Rule Faces September Deadline As Industry Seeks Narrower FilingsCapital & PolicyJul 26, 2026CIRCIA Rule Faces September Deadline As Industry Seeks Narrower FilingsCISA is working toward a September target for the delayed CIRCIA rule as industry groups seek fewer covered entities, narrower incident triggers and leaner reporting requirements. The law sets 72-hour incident and 24-hour ransomware-payment reporting deadlines, while the proposed rule could cover more than 300,000 entities.Ropedia Raises US$22 Million For Physical AI Data LayerAIJul 26, 2026Ropedia Raises US$22 Million For Physical AI Data LayerSingapore-based Ropedia has raised US$22 million in pre-Series A funding, e27 reported, as robotics developers look for data that records motion, depth, audio and human interaction rather than only online text.Azure West US Outage Exposes Fiber-Maintenance Routing RiskCloud & Data CentersJul 26, 2026Azure West US Outage Exposes Fiber-Maintenance Routing RiskThe Register reported that Microsoft’s West US Azure region was disrupted for almost five hours after routine fiber-related maintenance removed more routes than intended. Microsoft’s preliminary review said 27 services were affected before full recovery at 19:41 UTC on July 23.Coinbase Adds x402 Payments For AI Agent TransactionsFintech & Digital PaymentsJul 26, 2026Coinbase Adds x402 Payments For AI Agent TransactionsCoinbase Business users will be able to accept payments from AI agents through x402, while exchange users get live order views and developers get an SDK.Microsoft Azure AI Deals Add Cobalt Chips And French CapacityCloud & Data CentersJul 26, 2026Microsoft Azure AI Deals Add Cobalt Chips And French CapacityMicrosoft expanded Azure AI agreements with Databricks and Mistral, combining Cobalt processor adoption, Microsoft product integrations and Mistral-operated French data-centre capacity for customers with control and residency requirements.