![[Paper Review] Continuous Autoregressive Language Models](https://discuss.pytorch.kr/uploads/default/original/2X/d/d1107e24375fae20a5f3a4733826880ce5d817db.png)
[Paper Review] Continuous Autoregressive Language Models
Paper CALM: Bypassing the Token-by-Token Bottleneck with “Continuous Vector-by-Vector” Likelihood-Free Language ModelingCALM …
43 min
semantic-bandwidth
scaling-law
inference-efficiency
autoencoder
brierlm
![[Paper Review] Memory Retrieval and Consolidation in Large Language Models through Function Tokens](https://moonlight-paper-snapshot.s3.ap-northeast-2.amazonaws.com/arxiv/memory-retrieval-and-consolidation-in-large-language-models-through-function-tokens-4.png)
![[Paper Review] NVIDIA Nemotron 3: Efficient and Open Intelligence](https://developer-blogs.nvidia.com/wp-content/uploads/2025/12/image3-8-png.webp)
![[Paper Review] Radial Attention: O(n log n) Sparse Attention with Energy Decay for Long Video Generation](https://moonlight-paper-snapshot.s3.ap-northeast-2.amazonaws.com/arxiv/radial-attention-onlog-n-sparse-attention-with-energy-decay-for-long-video-generation-2.png)