klotz: prefill-decode disaggregation*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. We introduce Prefill-as-a-Service (PrfaaS), an architecture that enables LLM serving across distributed, cross-datacenter clusters. Unlike traditional disaggregation limited by RDMA requirements, PrfaaS leverages hybrid-attention models to reduce KV cache transfer needs. The system selectively offloads long-context prefill tasks to remote compute-dense clusters while handling short requests locally. Using a dual-timescale scheduler, PrfaaS optimizes routing and resource allocation over commodity Ethernet.


    - Overcomes bandwidth bottlenecks in disaggregated serving.
    - Uses hybrid attention to minimize KV cache throughput demands.
    - Implements selective offloading based on request length.
    - Employs dual-timescale scheduling for efficient traffic management.
    - Shows high throughput gains with 1T-parameter models.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: prefill-decode disaggregation

About - Propulsed by SemanticScuttle