Prediction Sets for Counterfactual Decisions: Coverage, Optimality, and Conformal Prediction
Studies uncertainty quantification via conformal prediction for counterfactual decision-making in high-stakes applications.
Studies uncertainty quantification via conformal prediction for counterfactual decision-making in high-stakes applications.
Addresses failure of on-policy self-distillation on long chain-of-thought reasoning, proposing method to maintain model thinking capability.
Proves aggregation with exponential weights is minimax-rate optimal in expectation for model selection, settling open problem from 2013.
Enables in-context learning in spiking neural networks via dendritic computation, making biologically plausible SNNs pass Garg-2022 ICL benchmark.
Redesigns symbolic parser backend using CCG directed types for improved structural generalization on SLOG benchmark with 30K parameters.
HNSW search framework adding theoretical correctness guarantees to hierarchical navigable small world graphs via graph spanner verification.
Proposes quantum sequence modeling using variational circuits with self-modulating gates and bounded memory for stable long-sequence processing.
Studies whether LLM personas from psychometric questionnaires are intrinsic or frame-dependent using geometric analysis on manifolds.
NASA deploys agentic search system using LLMs to help geoscience researchers discover relevant datasets and tools from thousands of available resources.
WattGPU predicts power consumption and latency for LLM inference across unseen GPUs without exhaustive profiling, addressing data center energy optimization.
arXiv paper on fast multi-dimensional refusal subspace extraction in LLMs for safety and interpretability.
arXiv paper on object-centric LeJEPA for more data-efficient self-supervised image representation learning.
Q-GAIN Python package for machine learning and physics-informed analysis of cold-atom experiment images with classification and detection.
OrbitQuant data-agnostic quantization method for diffusion transformers handling activation shifts across timesteps without recalibration.
CNeVA framework for controllable simulated traffic agents with interpretable behavior latents enabling edge case testing and variable isolation.
Study of how social structure and audience context affect what LLM agents express in multi-agent debate settings using dual-channel framework.
Real-time safety monitoring framework for LLMs using external verifiers with risk-calibrated thresholds to detect unsafe outputs at deployment.
LACUNA testbed evaluates parameter-level localization precision for LLM unlearning, addressing memorized sensitive training data removal.
MetaTT tensor-train adapter for parameter-efficient fine-tuning of transformers with flexible factorization across layers and task dimensions.
BALF framework for parameter-efficient model compression using activation-aware low-rank factorization beyond linear layers.
Method using LLM priors to enable efficient program learning through empirical risk minimization with fewer samples and less computation.
Framework incorporating latent geometry as explicit representation quality component under data scarcity through variational information bottleneck.
Theoretical analysis of deep neural network approximation rates for symmetric Korobov functions with polynomial dimension dependence.
Method for continual unlearning in diffusion models to progressively remove concepts while maintaining generation quality across multiple removal steps.
ThreadWeaver enables parallel reasoning in LLMs through adaptive threading to reduce inference latency while maintaining output quality.
ZENITH optimizer for automatic learning rate scheduling in deep vision models with lower computational overhead than existing adaptive optimizers.
Theoretical analysis comparing predictive inverse dynamics models to behavior cloning for offline imitation learning with limited demonstrations.
Research on spectral imbalance in low-rank continual learning for parameter-efficient model adaptation without catastrophic forgetting.
Efficient LLM deployment technique combining token-adaptive layer execution with quantization for reduced computation and memory.
Principled approach for upscaling smaller trained models to larger ones with hyperparameter transfer and warm starts.
Framework enabling language models to overcome context limitations by recursively invoking themselves to solve long-horizon reasoning problems.
Lightweight uncertainty quantification method for neural networks using gradient norms and isotropy assumptions.
Parameter-efficient LLM architecture using looped transformers to improve memory efficiency for edge and on-device deployment.
Federated fine-tuning framework using Fisher-guided token quantization to reduce communication for LLM adaptation on edge devices.
Geometric interpretation of transformer components showing attention and normalization emerge from polar state estimation.
Technique for KV caching shared prefixes in diffusion language models with bidirectional attention mechanisms.
Benchmark with 40 tasks across 10 scientific domains for evaluating end-to-end autonomous research capabilities of AI coding agents.
Framework for certifying when conservation laws remain valid in learned latent representations of physical systems.
Analysis of learning rate scaling laws for asynchronous RLHF with stale rollouts in high-throughput LLM training.
Study showing single transformer layer RL training matches full-parameter fine-tuning for LLM post-training with GRPO.
Introduction to Transformer architecture, key refinements, and applications in natural language processing.
Self-supervised learning approach inspired by neuroscience using predictive coding with biologically plausible credit assignment.
Survey of foundation models for VLSI circuit design and EDA using self-supervised pre-training on circuit data.
PPO-driven adaptive filtering with composite reward design for denoising in dynamic, non-stationary environments like wireless signals and biomedical monitoring.
Domain-adaptive continuous pre-training specializes LLMs for cybersecurity analysis with minimal tokens and HPC efficiency for reduced computational requirements.
Deep learning approach for cardiovascular disease detection via heart sound classification using synthetic and augmented phonocardiogram and electrocardiogram signals.
BuilderBench evaluates AI agents on acquiring skills through interaction and exploration rather than mimicry, measuring scalable learning mechanisms for novel problem-solving.
250m resolution NEXRAD radar dataset for machine learning precipitation nowcasting with fine-scale storm structures enabling extreme weather prediction.
Model merging technique navigates alignment-calibration trade-off by interpolating LLM weights, achieving Pareto-superior frontier without sacrificing task accuracy or calibration.
Computational framework using deep learning and LLMs to model human neurophysiological adaptation to altered gravity in spaceflight scenarios.