Paper Bites / Archive

Small notes.
Careful reading.

Each Paper Bite is a focused Markdown note: why the paper matters, the core idea, how it works, what to inspect in the results, and practical takeaways.

22research notes
and counting

Newest first
01

2026.10.09 / 7 min read / CVPR 2026

Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models

Beta-KD learns how strongly a multimodal student should follow its teacher, but its gains depend on the loss and comparison being examined.

Knowledge Distillation / Multimodal Models / Vision-Language Models
02

2026.10.06 / 7 min read / arXiv preprint

NovaLAD: A Fast, CPU-Optimized Document Extraction Pipeline for Generative AI and Data Intelligence

NovaLAD splits document parsing into parallel semantic and layout detection, then orders text and gates images before optional vision-model enrichment; its DP-Bench lead is author-reported against historical baseline rows.

Document Parsing / Layout Analysis / Optical Character Recognition
03

2026.09.29 / 6 min read / arXiv preprint (2025)

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

TurboQuant rotates vectors before scalar quantization and adds a residual sketch when unbiased inner products matter; its theory and two application tests need different readings.

Vector Quantization / KV Cache / Nearest Neighbor Search
04

2026.09.29 / 6 min read / arXiv preprint (2021)

SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval

SPLADE v2 improves a learned sparse retriever through max pooling, a document-only option, and distillation—but its strongest scores and efficiency claims belong to different configurations.

Information Retrieval / Neural Search / Sparse Retrieval
05

2026.09.07 / 6 min read / ICML 2024 (Oral)

DoRA: Weight-Decomposed Low-Rank Adaptation

DoRA keeps LoRA's low-inference-cost design but separates magnitude from direction, giving the adapter more room to change how a pretrained weight is scaled without updating the whole matrix.

Parameter-Efficient Fine-Tuning / LoRA / Large Language Models
06

2026.08.24 / 7 min read / arXiv preprint

Prime Agent: A Self-Improving RLM Harness

Prime Agent treats persistent computation, recursive subagents, and revisable harness state as part of the agent—not invisible plumbing—and evaluates how that substrate changes long-horizon work.

Agentic AI / Test-Time Compute / Multi-Agent Systems
07

2026.08.24 / 6 min read / arXiv preprint

GenRec: An LLM-Backed Recommendation Ranker at Netflix

GenRec verbalizes Netflix interaction histories for an adapted language model, then uses a catalog-aware scoring head and prefill-only inference to rank in-catalog items under production cost constraints.

Recommender Systems / Large Language Models / Learning to Rank
08

2026.08.24 / 6 min read / SIGIR 2020

ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT

ColBERT precomputes contextual token embeddings for documents, then scores each query through a cheap MaxSim late-interaction step that preserves fine-grained matching without running BERT on every query-document pair.

Information Retrieval / Neural Search / Dense Retrieval
09

2026.08.23 / 7 min read / NeurIPS 2020

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

The original RAG paper couples a trainable DPR query encoder with a fixed Wikipedia passage index and a BART generator, then marginalizes generation probabilities over retrieved latent documents.

Retrieval-Augmented Generation / Dense Retrieval / Open-Domain Question Answering
10

2026.08.14 / 5 min read / arXiv preprint (v2, 2025)

Small Language Models are the Future of Agentic AI

This position paper argues for SLM-first agents: route repetitive, narrow calls to specialized small models and retain large models for tasks that need broad generality.

Agentic AI / Small Language Models / Model Routing
11

2026.08.14 / 6 min read / SOSP 2023

Efficient Memory Management for Large Language Model Serving with PagedAttention

PagedAttention treats the key-value cache like virtual memory, letting vLLM allocate and share fixed-size blocks on demand so more requests fit into each serving batch.

LLM Serving / Systems / Memory Management
12

2026.07.07 / 7 min read / CVPR 2026

MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models

MMR-AD builds a large multimodal industrial anomaly benchmark with reasoning traces, localization annotations, and cross-dataset evaluation splits, then uses it to show that current generalist MLLMs still struggle badly on practical anomaly detection and especially anomaly localization.

Anomaly Detection / Multimodal Models / Datasets
13

2026.07.05 / 6 min read / CVPR 2026

MoECLIP: Patch-Specialized Experts for Zero-shot Anomaly Detection

MoECLIP adapts CLIP to zero-shot anomaly detection by routing each patch through specialized LoRA experts, then explicitly separating those experts so patch-level specialization improves without giving up CLIP's transfer ability.

Zero-Shot Anomaly Detection / CLIP / Mixture of Experts
14

2026.05.28 / 7 min read / NeurIPS 2025

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

ThinkSound treats video-to-audio as a reasoning problem rather than a one-shot mapping, using multimodal chain-of-thought to guide a unified model through foley generation, object-focused refinement, and instruction-based editing.

Audio Generation / Multimodal Models / Chain-of-Thought
15

2026.05.28 / 6 min read / CVPR 2023

DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation

DreamBooth turns subject-driven image generation into a lightweight personalization problem: bind a rare token to a specific subject, then regularize fine-tuning so the model keeps both the subject identity and the broader class prior.

Diffusion Models / Personalization / Image Generation
16

2026.05.26 / 6 min read / CVPR 2023

Simulated Annealing in Early Layers Leads to Better Generalization

SEAL improves iterative training by perturbing early layers with short bursts of gradient ascent instead of resetting later layers, leading to better in-distribution accuracy and much stronger transfer behavior than LLF.

Generalization / Optimization / Transfer Learning
17

2026.05.25 / 7 min read / arXiv 2026

Lance: Unified Multimodal Modeling by Multi-Task Synergy

Lance argues that a unified image-video model works better when tasks share context but not every representation or parameter path, combining multi-task training with separate understanding and generation routes.

Multimodal Models / Image Generation / Video Generation
18

2026.05.22 / 7 min read / arXiv 2025

Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset

Ditto turns instruction-based video editing into a data problem first: it builds a million-sample synthetic dataset, then trains Editto with curriculum learning to follow text edits more reliably.

Video Editing / Generative Models / Datasets
19

2025.05.28 / 7 min read / NeurIPS 2025

Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment

Chain-of-Zoom treats extreme super-resolution as a sequence of zoom steps, using multi-scale prompts and preference alignment to guide details beyond a model's usual scale range.

Super-Resolution / Vision-Language Models
20

2025.05.26 / 6 min read / ICLR 2022

LoRA: Low-Rank Adaptation of Large Language Models

LoRA freezes a pretrained language model and learns small low-rank update matrices, making task adaptation much cheaper to train, store, and switch.

Efficient Tuning / Large Language Models
21

2025.05.21 / 7 min read / ICCV 2023

Adding Conditional Control to Text-to-Image Diffusion Models

ControlNet adds spatial conditioning to pretrained text-to-image diffusion models while protecting the original model's generation quality through zero-initialized control branches.

Generative Models / Diffusion Models
22

2025.05.20 / 7 min read / CVPR 2024

ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions

ViT-CoMer keeps the flexibility of a plain Vision Transformer, then adds a parallel convolutional branch so dense prediction tasks can use richer multi-scale features.

Computer Vision / Dense Prediction