We introduce Prefill-as-a-Service (PrfaaS), an architecture that enables LLM serving across distributed, cross-datacenter clusters. Unlike traditional disaggregation limited by RDMA requirements, PrfaaS leverages hybrid-attention models to reduce KV cache transfer needs. The system selectively offloads long-context prefill tasks to remote compute-dense clusters while handling short requests locally. Using a dual-timescale scheduler, PrfaaS optimizes routing and resource allocation over commodity Ethernet.
- Overcomes bandwidth bottlenecks in disaggregated serving.
- Uses hybrid attention to minimize KV cache throughput demands.
- Implements selective offloading based on request length.
- Employs dual-timescale scheduling for efficient traffic management.
- Shows high throughput gains with 1T-parameter models.