StepAudio 3 Gen Technical Report
Paper StepAudio 3 Gen: Unifying Speech, Song, Music, and Sound Effects with a Single Discrete Autoregressive (RVQ) ModelTL;DRStepAudio 3 …
Paper StepAudio 3 Gen: Unifying Speech, Song, Music, and Sound Effects with a Single Discrete Autoregressive (RVQ) ModelTL;DRStepAudio 3 …
Paper SAS: A Simple Attention Sparsification That Learns Context Ranking Directly Without DistillationTL;DR — In post-training attention …
Paper Marigold V2: Reviving an Image-Editing Diffusion Transformer (DiT) as a Monocular Depth EstimatorTL;DRThis study fine-tunes an …
Paper SQS: Fusing Pruning and Quantization into One Bayesian Learning — Where Spike-and-Slab Meets GMMTL;DR — Doing pruning and low-bit …
Paper X-AuT: Pruning Audio Encoders in Speech LLMs with Behavioral Probes and Cross-Scale DistillationTL;DRX-AuT is a framework that …
Paper Deadline-Filling Prefill Chunks: SLOWeave — An Adaptive Chunking Scheduler for LLM Serving TL;DR — Instead of a fixed prefill chunk …
Paper SMELT: Scaling Laws for Compute-Matched MoE Loop TransformersTL;DRLoop transformers scale depth by executing a layer block repeatedly, …
![[Paper Review] Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality](https://icml.cc/media/PosterPDFs/ICML%202024/32613.png)
Paper Link Structured State Space Duality: Unifying SSMs and Attention with Mamba-2 for 2–8× AccelerationTL;DRStructured State-Space Duality …
![[Paper Review] Massive Activations in Large Language Models](https://eric-mingjie.github.io/massive-activations/assets/main_teaser_final.png)
Paper Link Massive Activations, Hidden Biases: A Reinterpretation of Self-Attention’s Secrets TL;DRJust 4–10 extreme scalar values (×10,000) …
Paper CODE I/O: From Code I/O + Natural-Language CoT to General-Purpose Reasoning — Lifting 7B-30B LLMs by +2 Points on Average with Data …