GEM: Geometric Entropy Mixing for Optimal LLM Data Curation
GEM reformulates LLM pre-training data curation as variational optimization on hypersphere to address ontological misalignment and embedding anisotropy in data mixing.
GEM reformulates LLM pre-training data curation as variational optimization on hypersphere to address ontological misalignment and embedding anisotropy in data mixing.
Pair-In, Pair-Out proposes latent multi-token prediction combining input-side compression and output-side efficiency to reduce LLM inference costs.
NRLB is a multi-agent framework for plain language summarization that adapts outputs for diverse reader groups with different linguistic and cognitive abilities.
Head-to-head comparison of Claude Code and Codex executing autonomous gravitational wave data analysis pipelines on shared infrastructure without human intervention.
SafeRx-Agent proposes a multi-agent framework combining LLMs with safety verification for medication recommendation, addressing explainability and traceability in clinical decision-making.
Studies compute allocation strategies in LLM-guided evolutionary search across depth-breadth tradeoff using multi-armed bandit approach.
Pocket-Dentist: efficient multimodal LLM for on-device dental image analysis optimizing for inference speed and privacy.
Efficient sparse coding approach for multi-vector retrieval replacing k-means clustering with single-stage processing.
QASM-Eval dataset trains LLMs on OpenQASM-3 quantum programming including error correction, timing, and pulse-level operations.
Unicorn framework for scalable multi-dataset time series forecasting balancing channel-independence and channel-dependency modeling.
Multi-model study of linear representations in LLM deceptive alignment using synthetic dishonesty as controlled testbed.
Proposes RBF networks as alternative LLM architecture without deep neural networks claiming improved explainability.
NumLeak measurement framework detecting memorized benchmark data in frontier LLMs via API probes and white-box analysis.
LongDS-Bench evaluates long-horizon agentic data analysis with 68 real-world Kaggle-based tasks testing context tracking across multi-turn interactions.
ML research extending calibration theory to probabilistic label ranking prediction tasks with structured output spaces.
Framework for evaluating LLM distillation beyond output matching using bounded behavioral indistinguishability formal measure.
VeriGate improves GRPO reasoning model training by adding step-level verifier feedback to address sparse supervision and credit assignment.
Unified theoretical framework analyzing gradient aggregation methods in multi-objective optimization with convergence rate analysis.
Neuro-symbolic learning framework integrating neural networks with differentiable optimization for incorporating domain knowledge as logical rules.
Distributed multi-agent reinforcement learning approach for constrained coordination problems with separable agent dynamics.
Model extraction attack against graph neural networks using explainability interfaces in GMLaaS platforms.
Theoretical framework for universal multiclass transductive online learning with unbounded label spaces.
Mechanistic interpretability study recovering explicit Zeta map algorithm on Dyck paths from trained transformer.
Graph-conditioned mixture of experts framework for spatio-temporal traffic forecasting with node-wise expert specialization.
Balanced benchmark and unlearning method for evaluating machine unlearning across causal and relational knowledge.
Study of representation collapse during sequential post-training of LLMs with measurement suite for hidden states and LoRA updates.
Analysis of AI-style alignment signatures in LLMs through measurement and localization of post-training effects on representations.
Study of long-term effects of data selection strategies in multi-stage LLM fine-tuning and model adaptability.
Procedural generation system for creating field-scale earth models and seismic data for full waveform inversion training.
Method for early prediction of human behavioral strategy from process traces for adaptive systems.
Theoretical perspective on diffusion models as information-withholding techniques and their advantages in data-scarce settings.
Study showing supervised training degrades visual cortex alignment in neural networks across multiple biologically plausible learning rules.
Benchmarking five uncertainty quantification methods for turbine gas temperature degradation prediction.
Causal Sensitivity Score metric to evaluate hidden capabilities in clinical LLMs and agents through counterfactual case mutations.
Method for assigning predictability scores to trajectory windows across different dynamics regimes using ordinal estimation.
Multi-task machine learning framework for turbine engine health management and remaining useful life prediction using real-world fleet data.
Research on improving relative representations in neural networks using learned anchors and whitened inner products to enable modular AI systems.
AMNESIA: first large-scale open-source benchmark for medical machine unlearning, enabling selective forgetting in medical LLMs.
TASER: training-time regularization framework using Langevin Stein operators to improve robustness under distribution shift and adversarial perturbations.
Fine-tuning method for adapting diffusion/flow models to optimize rewards while satisfying constraints for molecular design.
CSULoRA: safety-preserving low-rank adaptation method preventing adversarial fine-tuning from weakening LLM alignment.
Study showing diffusion models memorize prototypical examples rather than atypical ones. Analysis of training data memorization patterns.
LARK: method for selecting learnable reasoning trajectories for efficient knowledge distillation from teacher to student models.
Using FinBERT embeddings with Transformers for financial forecasting instead of scalar sentiment scores. Application of representation learning to finance.
Research on learning control-relevant representations via empowerment objective for unsupervised skill learning in reinforcement learning.
BOKBO: conformal abstention layer for vision-language-action policies providing distribution-free safety guarantees on executed-violation rate.
Categorical framework unifying decision-making theories (planning, RL, causal intervention, online learning) using universal constructions.
Systematic study of test-time compute strategies in vision-language models across seven models and six benchmarks using feature-based scoring and voting.
Prompted Policy Optimization (PromptPO): using LLMs as black-box policy optimizers for reinforcement learning tasks with executable policy generation.
Lossless compression techniques to reduce GPU memory bottlenecks during ML training and inference without accuracy loss.