M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
M4-RAG benchmark for multilingual multimodal retrieval-augmented generation spanning 42 languages and 56 countries with vision-language models.
M4-RAG benchmark for multilingual multimodal retrieval-augmented generation spanning 42 languages and 56 countries with vision-language models.
First public dataset for automatic essay scoring and feedback generation in Basque language with 3,200 expert-annotated essays at C1 proficiency level.
Novel non-autoregressive video generation framework using flow matching with interleaved frame insertion and denoising for variable-length output.
Test-time padding defense method for adversarial detection and robust adaptation in vision-language models like CLIP without requiring retraining.
Adaptive discrete video tokenizer using information-theoretic compression for efficient long video sequence processing with variable information density.
Hybrid neuro-symbolic approach combining neural networks with rule verification for zero-hallucination clinical ICD-10 coding with automated knowledge base expansion.
Multi-step actor-critic reinforcement learning with Lyapunov stability certificates for data-efficient exponentially stabilizing control in high-dimensional spaces.
Economic study analyzing how delegating pricing to LLMs can facilitate collusion through propensity and fidelity parameters in duopoly settings.
Benchmark introducing Teleo-Spatial Intelligence to measure vision-language models on physical dynamics and intent reasoning in open-world environments.
Time Puzzles benchmark for evaluating iterative temporal reasoning in LLMs with tool use like web search in constraint-based date inference tasks.
APEX-SWE benchmark assessing frontier AI models on real-world software engineering tasks including integration and debugging work.
Multi-task instruction tuning for Arabic audio LLMs combining generative tasks (ASR, summarization) and discriminative tasks (dialect, emotion recognition).
Benchmark evaluating 24 open-source language models on Sinhala across Unicode, Romanized, and mixed-script variations for low-resource language performance.
Memory-augmented diffusion model for consistent multi-turn video editing by preserving previously generated regions across sequential edits.
Analysis of which LLM layers handle multilingual tasks and methods to tune them for improved language control in non-English settings.
Meta-cognitive reinforcement learning framework enabling agents to reason about reliability of their own learning process under uncertainty.
Study detecting AI-generated content in academic peer reviews at ICLR and Nature Communications using trained detection models on historical data.
CALM: class-conditional sparse attention vectors for improving audio-language model performance on discriminative tasks like audio classification.
Analysis of variance in agentic evaluations showing single-run pass@1 scores are unreliable, with 2.2-6.0% variance across 60,000 trajectories.
Training approach for reasoning models on unverifiable data without external verifiers, removing dependency on high-quality human-annotated reasoning data.
Energy-aware reinforcement learning for robotic manipulation of articulated components in infrastructure maintenance operations.
LLM-enhanced rumor detection via virtual node edge prediction, capturing semantic flow across propagation paths in social networks.
Query-conditioned selector approach for soft context compression in RAG systems, improving scalability and reducing redundant retrievals.
Analysis of LLM agent caching failures and structured intent canonicalization method using few-shot learning to improve cache effectiveness.
LAVIDA: zero-shot video anomaly detection using multimodal LLMs without requiring real anomaly examples, leveraging context understanding.
PhysMem: memory framework enabling vision-language model robot planners to learn and remember physical properties through autonomous experience.
AngelSlim: comprehensive toolkit for large model compression including quantization, speculative decoding, token pruning, and distillation.
Simulator for heterogeneous and disaggregated LLM serving infrastructure, supporting diverse accelerators and distributed deployment scenarios.
Framework for hierarchical visual recognition using large multimodal models with taxonomy-aware representation alignment.
Symbolic machine learning approach for failure detection in chemical processes, prioritizing explainability and safety over neural methods.
Security framework for hierarchical autonomy evolution in AI agents, addressing vulnerabilities as LLM-driven agents become more autonomous.
Graph in-context operator networks for spatiotemporal prediction enabling neural networks to infer solution operators from contextual examples.
PREBA combines PCA-weighted retrieval-augmented generation with LLMs and Bayesian averaging for zero-shot surgical duration prediction.
Method for real-time model predictive control using implicit maximum likelihood estimation to accelerate diffusion-based trajectory planning.
HyCon introduces hyperbolic geometry-based control for text-to-image models to steer generation away from unsafe content using parallel transport.
PA3 method for aligning conversational LLM agents with complex business policies using chain-of-thought reasoning without full policy context.
Study of compute allocation strategies for reasoning-intensive retrieval in long-horizon agent tasks, balancing query expansion and re-ranking costs.
Research on detecting when language models actively conceal knowledge through classifier training, finding larger models better at deception.
Framework using vision-language models as online reward generators for robotic manipulation policy refinement through reinforcement learning.
Fast-WAM investigates whether World Action Models for embodied control require test-time future imagination or can execute actions without iterative planning.
MHPO method for stable reinforcement learning using hazard-aware policy optimization with modulated importance ratio control for GRPO-based frameworks.
Study of adversarial robustness of open-source vision-language models LLaVA and Qwen2.5-VL against gradient-based attacks in e-commerce environments.
Cross-domain few-shot learning approach using vision-language models like CLIP with improved interpretability for fine-grained visual recognition tasks.
CoVerRL addresses consensus trap in label-free reasoning by using generator-verifier co-evolution to improve LLM reasoning without ground-truth supervision.
Agent Control Protocol (ACP) is a formal specification for governance of autonomous agents in institutional environments, using cryptographic admission control to validate agent actions before execution.
Case study evaluating LLM-generated language lessons in Duolingo from student perspective, addressing gap in profession-specific content.
Nemotron-Cascade 2: 30B MoE open model with agentic capabilities, achieving IMO Gold Medal performance with compact 3B activated parameters.
FinTradeBench: financial reasoning benchmark for LLMs evaluating decision-making over heterogeneous signals from regulatory filings and price data.
Review of automatic collaboration analysis methods using task-oriented conversational data resources.
Scalable prompt routing method using fine-grained latent task discovery to select optimal LLM from candidate pools.