Coding-agents can replicate scientific machine learning papers
Coding agents replicate scientific ML paper claims autonomously. Framework validates agent ability to verify computational results from paper materials.
Coding agents replicate scientific ML paper claims autonomously. Framework validates agent ability to verify computational results from paper materials.
Investigates shortcut learning in legal judgment prediction using UK Employment Tribunal data. Domain-specific ML study with limited tech relevance.
Large Behavioral Model learns customer decision-making from retail transaction data. Domain-specific application with limited broader ML significance.
Universal approximation theorem for operators on Banach spaces using projection methods. Theoretical ML research with limited practical applications.
Inference-time intervention method for steering LLM behavior across multiple conflicting attributes without parameter updates. Addresses multi-attribute alignment challenges.
M4V uses Mamba architecture for efficient text-to-video generation with linear-time sequence modeling.
GrAInS enables gradient-based inference-time steering of LLMs and VLMs without weight updates.
Evaluation of RAG versus long-context prompting for clinical reasoning tasks over electronic health records.
REAL method for KV cache compression in long-context LLMs via retrieval-reasoning and attention analysis.
Weak-to-strong generalization for LLMs using contrastive learning with implicit rewards to improve robustness.
Unsupervised network anomaly detection using variational graph autoencoders without requiring labeled datasets.
Pipeline for speaker-attributed LLM persona modeling from civic deliberation recordings for controlled simulations.
Practical guide for generating synthetic data with differential privacy to address data scarcity and representation issues.
ReinforceGen system combines task decomposition, imitation learning, and RL fine-tuning for long-horizon robotic manipulation.
MORL approach using preference-conditioned policies to recover dense Pareto fronts by addressing early scalarization and advantage cancellation issues.
arXiv paper on deploying generative AI for 911 call-taker training to address staffing and training scalability challenges.
arXiv paper on LLMbda Calculus: formal framework for information flow control in LLM agents to defend against prompt injection.
arXiv paper on SWE-Milestone: benchmark for evaluating AI agents on continuous software evolution with temporal dependencies.
arXiv paper on ML for network attack classification and synthetic data generation using adversarial methods.
arXiv paper on SLIDERS: LLM-based systematic evidence synthesis and reconciliation for comprehensive document analysis.
arXiv paper on causal fairness in ML by tuning derivatives to handle bias in protected attributes.
arXiv paper on embodied multi-agent coordination using dialogue to align partially-observable world models.
arXiv paper on AnchorMoE: interpretable multivariate time series classification using mixture-of-experts routing.
arXiv paper on enabling LLMs to self-modify and consolidate memories for continual learning beyond in-context knowledge.
arXiv paper on PhysAssistBench: benchmark for evaluating LLMs assisting physicians through coordinated clinical knowledge, EHR interaction, and patient communication.
arXiv paper on ECHO: selective turn memory and pruning techniques for long-horizon language agents with bounded context windows.
arXiv paper evaluating 9 LLMs' ability to communicate probabilistic information in natural language with consistency and calibration assessment.
arXiv paper on workflow-level jailbreak construction in IDE coding agents, showing safety failures across multi-turn task decomposition.
arXiv paper on chain-of-thought distillation optimization for recommendation systems using student-aware techniques.
Large-scale evaluation of uncertainty estimation methods across 22 languages in LLMs for multi-choice question answering.
Test-time training framework for steering robot foundation models toward task variants using human video demonstrations without fine-tuning.
Method extracting Riemannian geometric structure from pre-trained language model embeddings to understand sentence classification geometry.
Foundation model for sleep analysis using hierarchical contrastive learning on multimodal biosignals from CNS and ANS.
Efficient zero-shot context extension method for LLMs using dynamic bifocal RoPE to handle long-context applications without retraining.
Framework interpreting knowledge distillation mechanisms in LLMs via interaction decomposition to understand why various KD methods succeed.
System combining LLM-guided Mixture-of-Experts with survival analysis for interpretable Alzheimer's disease risk prediction from neuroimaging.
Quantization technique adjusting scale asymmetry for few-bit integer precision to reduce clipping errors on outliers.
Training method for Mixture-of-Experts models reducing memory-access overhead on edge devices via differentiable routing consistency loss.
Flow matching method using optimal transport coupling for controlled generation of molecules with target properties.
System for optimizing expert placement in distributed Mixture-of-Experts model serving via online proactive placement strategy.
Benchmark library for federated continual learning with standardized evaluation protocol across datasets, task splits, and data distributions.
Dataset for studying data pricing mechanisms in data marketplaces, addressing valuation challenges for data products.
GPU kernel optimization for LLM inference acceleration using moderately sparse weight matrices, addressing performance gaps in sparse matrix multiplication.
Method using LLMs and vision-language-action models to improve exploration in reinforcement learning by generating diverse policy perturbations via natural language prompts.
Research on how linear representations emerge during neural network training, foundational for interpretability methods like linear probes and activation steering.
Study of adversarial attacks on LLM safety representations via activation-guided optimization and refusal geometry.
GNN approach that encodes missing data patterns alongside values for handling incomplete datasets.
Continuous batching framework for efficient diffusion LLM serving with block-grained scheduling to reduce latency.
Dynamic routing system selecting between LLMs and VLMs for time series reasoning based on modality characteristics.
Toolkit for systematic evaluation of algorithmic fairness across intersectional subgroups and modeling lifecycle stages.