LecturaAgents multi-agent framework for adaptive personalized AI-assisted learning with embodied teaching methods tailored to individual learners.
Studies primacy bias in multimodal retrieval-augmented QA; vision-language models show U-shaped 'lost-in-the-middle' effect when conditioning on retrieved context.
Interprets Transformer pretraining as fast-slow ODE dynamics with depth-ordered layer factorization, relating attention to coupled differential equations.
Unified taxonomy of distributional shifts in RL covering both in-distribution/out-of-distribution generalization and non-stationary environment dynamics.
OrthoReg combines symbolic physics-based models with neural networks for interpretable hybrid dynamical systems modeling using orthogonal regularization.
Analysis of pass@k difficulty metric reveals blind spot on hardest math reasoning problems; 10-23% of unsolvable problems misclassified as reachable in GSM8K/MATH.
Triangular consistency constraint for optical flow learning applicable across architectures and supervision types, enforcing cycle consistency in flows.
ELVA improves multimodal retrieval using ranking-driven contrastive learning with MLLMs to address grain-level information blindness in universal retrieval.
NRT-Bench benchmarks multi-turn adversarial attacks on LLM operator agents in safety-critical control room scenario with nuclear power plant simulation.
Digital twin framework using Unity for UAV pavement inspection under traffic conditions with procedurally generated defects and dynamic traffic simulation.
VideoAgent agentic framework for comprehensive video understanding and editing tasks, handling long-video coherence and diverse editing operations.
UC-Search framework for constrained time-series control using retained search to find feasible decisions with action constraints in delayed control settings.
Study showing LLM memory with incorrect cached conclusions causes worse performance than no memory, identifying 'brittle memory' failure mode in language models.
Privacy-preserving semantic search using SVD-truncation and homomorphic encryption to protect vector databases from inversion attacks while maintaining ranking quality.
Machine learning research on heavy-ball Q-learning method with convergence analysis and acceleration guarantees for reinforcement learning.
Research on LLM alignment techniques for code generation, examining trade-offs between functional correctness and non-functional code quality requirements.
Automated architecture search system that discovers optimal compositions of perception, memory, and planning modules for embodied agents.
Theoretical framework explaining grokking phenomena via stochastic-geometric analysis of neural network solution space topology.
Dataset and method for detecting emotional entrainment in dyadic speech conversations with conversational AI agents.
Benchmark for evaluating long-horizon stability of interactive world models across action, visual fidelity, memory, and physics dimensions.
Technical report on GR2, an LLM-based re-ranking system for industrial recommendation pipelines addressing deployment gaps.
Credit assignment method for agentic reinforcement learning that differentiates between exploration and regressive actions in multi-step agent trajectories.
LLM-based hard negative sampling technique for training two-tower retrieval models in large-scale recommendation systems.
Studies citation hallucinations in LLM-generated scientific papers submitted to peer-reviewed conferences, measuring prevalence and persistence.
Develops post-training pruning methods for diffusion transformers to reduce computational overhead in image generation models.
Examines how software engineering shifts toward AI code generation, focusing on maintainability, inspectability, and feedback loops for agentic development systems.
Block-diffusion decoder for faster generative reasoning re-rankers in recommendation systems, parallelizing token generation to reduce inference latency.
KV cache compression technique for accelerating long chain-of-thought reasoning in LLMs using sliding-window approach to reduce memory and latency.
Likelihood-based framework for automatic evaluation of turn-taking naturalness in full-duplex spoken dialogue systems using causal modeling.
Training-free attribution method for multimodal QA systems that maps model outputs to evidence using attention mechanisms for improved transparency in AI assistants.
Q-learning framework for dynamically selecting dialogue policies in high-stakes persuasion scenarios with LLMs, addressing personalized strategy selection for resident evacuation.
Method for predicting closed-loop performance of latent world models using validation-time diagnostics for offline checkpoint selection in MPC/RL.
kNNGuard: training-free guardrail using LLM hidden activations and k-NN to detect unsafe/off-topic prompts without fine-tuning.
Guided Action Flow: inference-time framework using learned action-chunk critic to guide frozen flow-matching vision-language-action policies.
Adaptive Reparameterized Time (ART) continuous-time control formulation learning optimal timestep allocation for score-based diffusion sampling.
Systematic study of AI agents patching compiler missed optimizations, identifying generalization challenges beyond fixing reported cases.
DemoPSD: disagreement-modulated policy self-distillation for LLMs reducing overfitting from privileged information while improving reasoning.
Analysis of five failure modes in benchmark-validity audits showing how implementation details can silently manufacture audit conclusions.
QuantFlow: federated Mamba-based foundation model for time-series forecasting with sequence embedding for privacy-sensitive signals.
GRAFT: per-word pronunciation conditioning mechanism for text-to-speech neural codec language models handling rare words and technical terms.
Federated learning approach for object detection enabling collaborative drone learning without centralizing aerial imagery data.
Post-hoc curation method for synthetic images via homogeneous-heterogeneous splitting to improve training data quality without retraining generators.
Research on training risk-averse language models to generalize risk aversion from low-stakes to high-stakes scenarios as AI safety approach.
Lagrangian Reward Augmentation (LARA) framework for inference-time alignment of frozen LLMs using auxiliary reward signals with explicit safety constraints.
Study of induction heads in transformers, identifying attention circuits that implement soft context-matching estimators for in-context learning via smoothing mechanisms.
Hybrid block diffusion language model training using partial bidirectionality to improve long-context generation efficiency and memory bandwidth.
Study of discrete diffusion model adaptation for molecular optimization with limited oracle budget and test-time task-specific optimization.
Data-free meta-learning method using pre-trained models to generate training tasks without access to labeled data.
Uncertainty estimation approach for RL-based algorithmic trading agents to handle market volatility and regime shifts.
Study comparing process-level vs outcome-only reward structures in reinforcement learning for mathematical reasoning in small language models.