Conditioning LLMs to Generate Code-Switched Text
Framework for conditioning LLMs to generate code-switched text focusing on English-Spanish language mixing capabilities.
Framework for conditioning LLMs to generate code-switched text focusing on English-Spanish language mixing capabilities.
Generative control policies using flow matching for robot tasks, addressing expert demonstration requirements and dynamic task execution.
FragFM framework for efficient molecular generation using fragment-level discrete flow matching with hierarchical autoencoder.
System-level DPO method for aligning compound AI systems with multiple interacting components via policy optimization.
FindAnything open-world mapping framework using vision-language models for real-time semantic understanding in robot exploration.
Controlled study analyzing tokenizer bias and LLM backbone capabilities for time series forecasting tasks.
Survey of federated learning as distributed ML paradigm enabling collaborative training across clients while preserving data privacy.
HCT-QA benchmark for question answering on human-centric tables in PDFs and web pages with complex structural layouts.
RM-R1 integrates chain-of-thought reasoning into reward models for improved LLM alignment with human preferences via RLHF.
Survey of benchmarks for code LLMs and agents across software development lifecycle stages with tiered taxonomy framework.
Event-based neural networks optimized for asynchronous event camera data using improved asynchronous-to-synchronous encoding methods.
AdAEM proposes automated measurement of LLMs' value differences and biases through adaptive evaluation methodology.
KramaBench benchmark evaluates AI systems on end-to-end data-to-insight pipelines over real-world data lakes with unstructured data.
SPARC introduces concept-aligned sparse autoencoders enabling cross-model and cross-modal interpretability for AI systems.
MAP mitigates hallucinations in vision-language models using map-level attention on hidden state semantics to improve factual consistency.
VLMQ applies token saliency-driven post-training quantization to vision-language models for efficient compression without retraining.
Geometric analysis using graph Ricci curvature to explain why GNN-based SAT solvers degrade on harder constraint problems.
Framework for assessing performance of language models in healthcare applications, accounting for variability in real clinical environments.
Answer-Then-Check safety alignment method enhances LLM robustness against jailbreak attacks by applying reasoning before generating final responses.
LikePhys evaluates physical plausibility understanding in video diffusion models, separating physics correctness from visual appearance.
Phys2Real combines vision-language models with reinforcement learning for robust sim-to-real transfer in robotic manipulation tasks.
CanvasMAR improves masked autoregressive video generation by adding structured global priors to reduce distortion in frame sampling.
Framework that infers user objectives in real-time to specialize LLM outputs and create customized tools, interfaces, and responses on-the-fly.
Method for 3D spatial reasoning in vision-language models, improving performance on tasks requiring understanding of spatial relationships from limited views.
Evaluates ChatGPT's consistency in coding communication data across different demographic subgroups, assessing reliability for large-scale analysis tasks.
Methods to evaluate and improve AI agents' information-seeking behavior under uncertainty, drawing from human cognition research for high-stakes applications.
LA-MARRVEL is an LLM framework for clinical rare disease diagnosis that matches genes to patient phenotypes, improving diagnostic speed by 12-15 percentage points.
Study of cultural memory and memorization in text-to-image diffusion models through multimodal iconicity phenomenon.
SQDF: KL-regularized RL method for diffusion model fine-tuning via soft Q-function to mitigate reward over-optimization.
Analysis of reverse KL divergence in RL-tuned LLMs explaining diversity loss, proposing filtering-based reasoning improvements.
A-3PO: accelerated asynchronous LLM training with PPO using staleness-aware proximal policy approximation.
Analysis and improvements for hyperbolic deep reinforcement learning agents, identifying gradient factors affecting optimization success.
Tools Orchestration Privacy Risk: systematic study of how multi-tool agents leak sensitive information through tool aggregation.
CASA: efficient vision-language fusion using cross-attention over self-attention to reduce memory and compute for multi-image conversations.
CARE: failure-centric reinforcement learning framework for multimodal reasoning with contrastive learning from incorrect rollouts.
LLMTM: benchmark for evaluating LLMs on temporal motif analysis in dynamic graphs with optimization techniques.
WBC: membership inference attack against fine-tuned LLMs using window-based localized memorization signals instead of global loss.
Framework for fine-tuning LLMs to generate grade-appropriate educational content across elementary to adult education levels.
Case studies demonstrating how Google's Gemini LLMs assist researchers with mathematical discovery and routine scientific tasks.
Aletheia: an autonomous AI agent for mathematics research that generates, verifies, and revises proofs at professional research level.
Research on AI-assisted code generation via natural language instructions, examining human collaboration and productivity impacts.
DataChef uses reinforcement learning to automatically optimize data recipe design for LLM adaptation, treating data curation pipeline design as learnable process.
SWE-MiniSandbox enables scalable RL training of software engineering agents without containers, reducing storage and setup overhead.
Peak+Accumulation formula detects multi-turn prompt injection attacks by aggregating per-turn risk scores at proxy layer without invoking LLM.
IntelliAsk uses RLVR to train LLMs to generate high-quality research review questions by learning from expert annotations across effort, evidence, and grounding dimensions.
FLoRG enables federated fine-tuning of LLMs using low-rank adaptation with Procrustes alignment for distributed collaborative training without data sharing.
Research evaluates when speech LLMs behave identically to ASR→LLM cascades using matched-backbone testing and mechanistic analysis.
EMPO² framework combines on/off-policy RL optimization with memory-augmentation to improve exploration in LLM agents, addressing key bottleneck in novel state discovery.
Information-theoretic analysis of multimodal LLM failures, framing inference as mismatched decoding that ignores non-text-aligned information.
CoME: Mobile agent architecture with four experts for decoupled yet integrated hybrid-capabilities reasoning including planning and action decision.