HBF for High-Throughput LLM Serving: UC Berkeley and FuriosaAI's

HBF for High-Throughput LLM Serving (UC Berkeley, FuriosaAI)


As large language models (LLMs) continue to scale in size and complexity, the demand for high-throughput serving has become a critical bottleneck. In 2026, researchers from UC Berkeley and FuriosaAI are pioneering a novel approach using Hybrid Bonding Foundation (HBF) technology to dramatically improve LLM inference performance.


What is HBF?


HBF, or Hybrid Bonding Foundation, is an advanced packaging technique that enables dense, high-bandwidth interconnects between logic and memory dies. By stacking memory directly on top of compute using hybrid bonding, HBF achieves significantly higher memory bandwidth and lower latency compared to traditional 2.5D or 3D stacking methods. This makes it particularly well-suited for memory-bound workloads like LLM serving.


Why HBF for LLM Serving?


LLM inference is heavily memory-bound: generating each token requires fetching billions of model parameters from memory. Traditional GPU-based serving is often limited by HBM bandwidth and capacity. HBF addresses these limitations by:


  • Increasing bandwidth density: Hybrid bonding allows for finer-pitch interconnects, enabling terabytes per second of memory bandwidth.
  • Reducing energy per bit: Shorter interconnect distances lower power consumption, crucial for large-scale deployments.
  • Enabling near-memory computing: HBF facilitates integrating compute elements closer to memory, reducing data movement.

UC Berkeley and FuriosaAI Collaboration


UC Berkeley's research focuses on system-level optimizations for HBF-based LLM serving, including memory management, scheduling, and compiler techniques. FuriosaAI, a leader in AI accelerators, contributes its expertise in high-performance, energy-efficient chip design. Together, they are developing a full-stack solution that leverages HBF to deliver unprecedented throughput for LLM inference.


Key Innovations


  • Unified memory hierarchy: A seamless memory space across HBF stacks simplifies programming and maximizes utilization.
  • Dynamic batching and paging: Advanced techniques to handle variable-length sequences and maximize hardware efficiency.
  • Hardware-software co-design: Joint optimization of the accelerator architecture and the serving stack.

2026 Context: The Year of HBF Adoption


In 2026, HBF is transitioning from research labs to commercial deployment. Major cloud providers and AI chip startups are racing to integrate HBF into their next-generation AI accelerators. The collaboration between UC Berkeley and FuriosaAI is at the forefront of this trend, demonstrating real-world viability for high-throughput LLM serving.


With LLMs becoming ubiquitous in applications ranging from chatbots to code generation, HBF-based serving promises to reduce costs, lower latency, and enable new capabilities. As the technology matures, we can expect HBF to become a standard building block for AI infrastructure.




Stay tuned for more updates on HBF and LLM serving from UC Berkeley and FuriosaAI.

via Semiconductor Engineering

Related