Software Framework Optimization for Long-Sequence LLM Inference
Developed the Multi-Granularity Bit-Level Hierarchical framework to optimize long-sequence LLM inference through bit-level aggregation and dynamic memory management.
- 1.45x latency reduction over dense attention
- 2.89x speedup over sparse-band baselines
- Minimal accuracy loss in BERT-based experiments