PMF-CL: Pareto-Minimal-Forgetting Continual Learner for Conflicting Tasks
PMF-CL algorithm for continual learning using Pareto-minimal-forgetting approach to minimize catastrophic forgetting on conflicting tasks.
PMF-CL algorithm for continual learning using Pareto-minimal-forgetting approach to minimize catastrophic forgetting on conflicting tasks.
Importance smoothing technique for training deep state space models at scale, combining variational and backpropagation approaches.
PRISM data selection method using preference-aware influence functions to select high-value examples for efficient LLM fine-tuning.
Agent JIT compilation method for web agents that reduces latency by batching multiple tool calls before LLM evaluation.
Game-theoretic framework analyzing trade-off between model utility and distillation attacks with adaptive student and teacher defenses.
Learned Relay Representations method for masked diffusion models to preserve internal computation between denoising iterations.
GEM framework for LLM pre-training data curation using geometric entropy mixing on hypersphere to optimize data composition.
Framework for structured classification using loss functions that exploit class relationships. Addresses structure-agnostic loss limitations.
Black-box method to infer lower bounds on LLM parameter counts from text memorization patterns without model access. Novel inference technique.
Dimension reduction framework for Bayesian inference in high-dimensional PDE inverse problems using variational flows. Specialized mathematical methods.
Framework for multilingual proof data synthesis and theorem proving. Infrastructure for interfacing theorem provers at repository scale with parallel execution.
Research showing chain-of-thought reasoning in LLMs exhibits unfaithfulness on natural prompts without explicit biases, revealing reasoning discrepancies.
Research on LoRA interference in model merging using orthogonal subspaces. Addresses performance degradation when combining multiple fine-tuned LMs.
Eso-LMs: Family of diffusion language models interpolating autoregressive and masked diffusion paradigms with parallel generation and KV caching.
Position paper arguing quantum kernel machines should explore matrix-valued kernels beyond scalar kernels for advantages.
DISCO method uses conditional distance correlation to mitigate dataset bias in deep learning through causal framework.
Conformal C2ST method improves two-sample testing by transforming weak classifiers into strong statistical tests.
Benchmarking uncertainty quantification methods in chest X-ray classification for trustworthy medical AI deployment.
Theoretical analysis of stochastic gradient optimization with unknown nuisance parameters, establishing convergence guarantees.
Neuro-symbolic approach integrating deep learning with temporal logic for business process monitoring and suffix prediction.
MatchFixAgent: Language-agnostic AI agent for autonomous code translation validation and repair across programming languages.
Research on audio classification showing global pooling creates bottleneck in frozen model probing; proposes using patch tokens instead.
arXiv benchmark for tabular anomaly detection with semantic context and domain knowledge metadata.
arXiv research on adversarial robustness in one-stage learning-to-defer systems for hybrid predictor-expert decision-making.
arXiv paper on unified spatio-temporal video understanding combining object detection, tracking, and natural language captioning.
arXiv paper on probabilistic robustness evaluation for deep learning models under unknown perturbation distributions.
arXiv research on instrumental variable regression using outcome-aware spectral features for causal effect estimation.
arXiv research on evaluating conditional coverage in conformal prediction methods for reliable predictive systems.
arXiv paper on domain-specific foundation models for agentic physical AI, case study on nuclear reactor control with safety constraints.
arXiv research on federated causal discovery handling heterogeneous client models with unknown interventions.
arXiv paper on few-shot 3D point cloud segmentation using decoupled expert approach for multimodal data.
arXiv study measuring authority bias in language models across mathematical, legal, and medical reasoning tasks with endorsement sources.
arXiv research on optimizing neural network layer sizes for speech models, addressing performance-complexity trade-offs during training.
SERA framework for efficiently training open-weight repository-specific coding agents with soft verification without expensive full training.
Token Sparse Attention method for efficient long-context LLM inference using dynamic interleaved token selection per layer/head.
BAT uses convex gated probing mechanism to evaluate audio SSL embeddings better than finetuning on AudioSet.
Theoretical analysis of error propagation and model collapse when training diffusion models recursively on synthetic data.
IAPO framework uses information-theoretic token-wise rewards for efficient reasoning in LLMs, reducing inference token costs.
KernelCraft benchmarks agentic LLM systems for generating low-level kernels on emerging AI accelerators with novel ISAs.
GradMem method using test-time gradient descent for compressive memory conditioning in LLMs, reducing KV-cache memory overhead.
Learning-to-defer framework allowing dynamic selection of expert-specific information like retrieved documents and tool outputs.
Method decomposing CLIP projectors to improve intra-modal alignment for image-to-image and text-to-text retrieval tasks.
Benchmark and toolkit for evaluating LLM agents on physics simulations with cost-aware metrics accounting for tool-use expenses.
Learning-to-defer system framework for routing inputs between classifier and multiple experts with theoretical improvements.
DeepInsight method for informal theorem proving with LLMs by identifying core solution techniques through insight recognition.
Utility-Aligned Embeddings framework combining dense retrieval with LLM utility for improved RAG performance and computational efficiency.
Research on uncertainty quantification for online optimization methods using sketching to reduce computational complexity.
Method for improving vision-language model robustness by exploiting multimodal redundancies to reduce hallucination on corrupted inputs.
Online learning-to-defer algorithm for routing queries between models and varying expert pools with bandit feedback.
ASH: agentic system for embodied policy learning from unlabeled internet video using self-improvement loop with inverse dynamics models.