StepAudio 3 Gen Technical Report
Paper StepAudio 3 Gen: Unifying Speech, Song, Music, and Sound Effects with a Single Discrete Autoregressive (RVQ) ModelTL;DRStepAudio 3 …
Paper StepAudio 3 Gen: Unifying Speech, Song, Music, and Sound Effects with a Single Discrete Autoregressive (RVQ) ModelTL;DRStepAudio 3 …
Paper SAS: A Simple Attention Sparsification That Learns Context Ranking Directly Without DistillationTL;DR — In post-training attention …
Paper SQS: Fusing Pruning and Quantization into One Bayesian Learning — Where Spike-and-Slab Meets GMMTL;DR — Doing pruning and low-bit …
Paper X-AuT: Pruning Audio Encoders in Speech LLMs with Behavioral Probes and Cross-Scale DistillationTL;DRX-AuT is a framework that …
Paper Why Are Video LLMs Still So Expensive? — A Four-Stage Guide to Inference Efficiency MechanismsTL;DR — This paper classifies 125 …
Paper Miles v0.1: A “verified, clean, and scalable” full-stack system for frontier RL post-trainingTL;DR: Miles is a full-stack …
Paper Link Peri-LayerNorm: A Third Option Beyond Post-LN and Pre-LNTL;DRBy simply adding another LayerNorm right after the residual …
Paper Janus-Pro 7B: Dual-Encoder Multimodal LLM That Outsmarts Bigger ModelsOne-line summary (TL;DR)By fully separating the SigLIP …