Paper Bites / Archive
Small notes.
Careful reading.
Each Paper Bite is a focused Markdown note: why the paper matters, the core idea, how it works, what to inspect in the results, and practical takeaways.
22research notes
and counting
All Paper Bites
Newest firstUncertainty-Aware Knowledge Distillation for Multimodal Large Language Models
Beta-KD learns how strongly a multimodal student should follow its teacher, but its gains depend on the loss and comparison being examined.
NovaLAD: A Fast, CPU-Optimized Document Extraction Pipeline for Generative AI and Data Intelligence
NovaLAD splits document parsing into parallel semantic and layout detection, then orders text and gates images before optional vision-model enrichment; its DP-Bench lead is author-reported against historical baseline rows.
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
TurboQuant rotates vectors before scalar quantization and adds a residual sketch when unbiased inner products matter; its theory and two application tests need different readings.
SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval
SPLADE v2 improves a learned sparse retriever through max pooling, a document-only option, and distillation—but its strongest scores and efficiency claims belong to different configurations.
DoRA: Weight-Decomposed Low-Rank Adaptation
DoRA keeps LoRA's low-inference-cost design but separates magnitude from direction, giving the adapter more room to change how a pretrained weight is scaled without updating the whole matrix.
Prime Agent: A Self-Improving RLM Harness
Prime Agent treats persistent computation, recursive subagents, and revisable harness state as part of the agent—not invisible plumbing—and evaluates how that substrate changes long-horizon work.
GenRec: An LLM-Backed Recommendation Ranker at Netflix
GenRec verbalizes Netflix interaction histories for an adapted language model, then uses a catalog-aware scoring head and prefill-only inference to rank in-catalog items under production cost constraints.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
ColBERT precomputes contextual token embeddings for documents, then scores each query through a cheap MaxSim late-interaction step that preserves fine-grained matching without running BERT on every query-document pair.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
The original RAG paper couples a trainable DPR query encoder with a fixed Wikipedia passage index and a BART generator, then marginalizes generation probabilities over retrieved latent documents.
Small Language Models are the Future of Agentic AI
This position paper argues for SLM-first agents: route repetitive, narrow calls to specialized small models and retain large models for tasks that need broad generality.
Efficient Memory Management for Large Language Model Serving with PagedAttention
PagedAttention treats the key-value cache like virtual memory, letting vLLM allocate and share fixed-size blocks on demand so more requests fit into each serving batch.
MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
MMR-AD builds a large multimodal industrial anomaly benchmark with reasoning traces, localization annotations, and cross-dataset evaluation splits, then uses it to show that current generalist MLLMs still struggle badly on practical anomaly detection and especially anomaly localization.
MoECLIP: Patch-Specialized Experts for Zero-shot Anomaly Detection
MoECLIP adapts CLIP to zero-shot anomaly detection by routing each patch through specialized LoRA experts, then explicitly separating those experts so patch-level specialization improves without giving up CLIP's transfer ability.
ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
ThinkSound treats video-to-audio as a reasoning problem rather than a one-shot mapping, using multimodal chain-of-thought to guide a unified model through foley generation, object-focused refinement, and instruction-based editing.
DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
DreamBooth turns subject-driven image generation into a lightweight personalization problem: bind a rare token to a specific subject, then regularize fine-tuning so the model keeps both the subject identity and the broader class prior.
Simulated Annealing in Early Layers Leads to Better Generalization
SEAL improves iterative training by perturbing early layers with short bursts of gradient ascent instead of resetting later layers, leading to better in-distribution accuracy and much stronger transfer behavior than LLF.
Lance: Unified Multimodal Modeling by Multi-Task Synergy
Lance argues that a unified image-video model works better when tasks share context but not every representation or parameter path, combining multi-task training with separate understanding and generation routes.
Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
Ditto turns instruction-based video editing into a data problem first: it builds a million-sample synthetic dataset, then trains Editto with curriculum learning to follow text edits more reliably.
Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment
Chain-of-Zoom treats extreme super-resolution as a sequence of zoom steps, using multi-scale prompts and preference alignment to guide details beyond a model's usual scale range.
LoRA: Low-Rank Adaptation of Large Language Models
LoRA freezes a pretrained language model and learns small low-rank update matrices, making task adaptation much cheaper to train, store, and switch.
Adding Conditional Control to Text-to-Image Diffusion Models
ControlNet adds spatial conditioning to pretrained text-to-image diffusion models while protecting the original model's generation quality through zero-initialized control branches.
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions
ViT-CoMer keeps the flexibility of a plain Vision Transformer, then adds a parallel convolutional branch so dense prediction tasks can use richer multi-scale features.