Calibrated Test-Time Guidance for Bayesian Inference
Analysis of test-time guidance in diffusion models showing common methods miscalibrate Bayesian inference; proposes correction for posterior sampling.
Analysis of test-time guidance in diffusion models showing common methods miscalibrate Bayesian inference; proposes correction for posterior sampling.
MetaOthello controlled study examining how transformers organize multiple world models across different generative processes using Othello game variants.
Amortized maximum inner product search: neural networks trained to directly predict MIPS solutions, amortizing costs for repeated queries on fixed databases.
Self-improvement framework maximizing mutual information between prompts and LLM responses without additional labeled data or external verifiers.
ActivityNarrated dataset for open-ended narrative-based human activity recognition from wearables, replacing fixed-window classification benchmarks.
Crystalite: lightweight diffusion transformer for crystal material modeling using subatomic tokenization and equivariant inductive biases.
Orthogonal BackFill method for compressing KV-cache communication in multi-agent LLM systems, reducing memory and communication costs while preserving information.
GUI-Perturbed framework reveals brittleness in GUI grounding models through domain randomization, showing 27-56% accuracy drops on spatial reasoning tasks.
Zeroth-order optimization methods for gradient-free black-box learning and memory-efficient LLM fine-tuning, analyzing stability dynamics.
REALM method for fine-tuning LLMs with noisy crowdsourced annotations by jointly learning model parameters and annotator expertise weights.
Fisher-guided token quantization reduces communication overhead in federated fine-tuning of LLMs on edge devices.
FastUMAP landmark-based dimensionality reduction method optimized for repeated analysis with changing hyperparameters.
POMDP approach with spatio-temporal attention for federated learning client selection under partial visibility constraints.
Meta-feature analysis explains performance gaps between tabular foundation models and traditional models on prediction tasks.
Protocol for diagnosing false positives in tail-aware LLM evaluation metrics beyond mean-based assessment.
Regime-stratified evaluation reveals hidden failures in time series foundation models masked by standard aggregate metrics.
Neuron-wise sequence modeling framework allowing independent evolution of neurons instead of layer-wise shared dynamics.
PersistentKV optimizes decode scheduling for long-context LLM serving on GPUs through page-aware KV cache management.
Closed-form reduced-order model of GRPO training dynamics using mean-field approximation and stochastically-forced oscillator analogy.
Fora protects LLM capabilities during fine-tuning by preserving activation subspaces rather than just parameter distances.
ComplianceGate routes LLM queries through multi-tier classifier system for compliance enforcement and cost efficiency in regulated industries.
Analysis of barren plateaus in quantum machine learning using Lie algebra perspective to address expressivity-trainability tradeoffs.
Pattern Embedded Neural Networks (PENNs) for multivariate regression with missing covariates, combining imputation with indicator observation networks.
Efficient sparse attention kernel implementation for LLMs, providing alternative to Native Sparse Attention with improved performance on long-context tasks.
Test adequacy metric (Clotho) for measuring LLM input quality before inference to enable task-specific prompt testing without ground truth outputs.
Large-scale dataset of 400K visual chain-of-thought reasoning examples for training multimodal large language models on spatially grounded reasoning tasks.
Dynamic sparse attention mechanism for video diffusion transformers to reduce quadratic complexity of self-attention in long-sequence generation tasks.
MMLoP low-rank prompting method for efficient vision-language model adaptation with millions fewer parameters than state-of-the-art.
DICE-RL framework using RL as distribution contraction operator to finetune pretrained diffusion/flow-based robot policies from online feedback.
IRIS benchmark: 240 4K real-world videos for unsupervised physical parameter estimation and governing-equation identification.
Surrogate modeling framework for stochastic differential equations using path-space observable error bounds for efficient simulation.
LLM-powered pipeline for automated extraction and structuring of materials science data from unstructured scientific literature.
Neural network emulators for real-time tokamak plasma shape control via virtual circuits; replaces expensive numerical computation.
Physics-informed review of deep learning statistical properties including neural scaling laws from classical statistics perspective.
FeLoG distributed graph embedding framework with feedback loop mechanism for billion-scale graphs; applications include GraphRAG.
Taxonomic strategy retrieval approach to mitigate compounding failures in multi-step agentic tasks; addresses problem drift in persuasion agents.
DataComp-VLM benchmark with 160 datasets for systematically evaluating vision-language model training data curation strategies.
Study of evaluation inconsistencies in diffusion LLM decoding strategies; identifies reproducible sources of bias in efficiency benchmarking.
HASTE hierarchical multi-agent system for ML engineering that accumulates cross-competition knowledge via LLM-driven abstraction across three scope tiers.
Academic perspective on AI's role in scientific research infrastructure and risks of homogenization.
Open-source auditable sandbox for recording and monitoring AI coding agent actions, behaviors, and file modifications.
UI design tool for real devices with export of layout specifications and prompts for AI code generation.
HN discussion thread asking about local LLM deployments in organizations, hardware choices, and operational challenges.
Open-source DOCX editor and SDK for rendering, editing, and automating Word documents in browsers and AI agent workflows.
Founder perspectives on scaling AI adoption across organizations and engineering teams.
DocETL enables processing large data collections with LLMs using natural language operations and map-reduce patterns.
Open-source Windows desktop app for AI workspace. Supports cloud APIs, local LLMs, tool agents, and workflow orchestration with governance tracking.
Case studies on using TimescaleDB open-source database for energy and carbon data analysis at scale.
Open-source duplicate code detector for AI-generated code using structural fingerprinting instead of text matching.
Gnosys: autonomous model engineer improving prompts and classifiers under label scarcity using optimization techniques.