The Bystander Effect in Multi-Agent Reasoning: Quantifying Cognitive Loafing in Collaborative Interactions
Study of cognitive loafing in multi-agent LLM systems showing collaboration can degrade reasoning quality due to social pressure effects.
Study of cognitive loafing in multi-agent LLM systems showing collaboration can degrade reasoning quality due to social pressure effects.
Research on limitations of cross-lingual transfer for low-resource NLP using Luxembourgish as case study, showing need for language-specific approaches.
AURORA micro-agent framework for diagnosing and mitigating grey failures in edge computing with causal reasoning under uncertainty.
Visual question answering dataset from satellite imagery for spatiotemporal analysis of construction activity.
One-shot federated learning framework with token relabeling for vision transformers under non-IID data distributions.
Adaptive frame selection method for efficient long-video understanding in vision-language models via posterior probing.
Entropy maximization approach for untargeted jailbreaks on vision-language models with improved transferability.
Cross-modal prompt generation framework for multimodal continual instruction tuning to mitigate catastrophic forgetting.
Dynamic mixture-of-experts MLLM for remote sensing scene segmentation with multimodal caption guidance.
Application of large language-vision models to remote sensing automatic target recognition tasks.
Visual tokenization method using multi-layer representation fusion from frozen vision encoders for better reconstruction.
Study of involuntary information leakage from prompted secrets in language model outputs using multi-model detection.
Identifies format confound in chain-of-thought corruption studies that detects answer text location rather than actual computation.
Benchmark suite for evaluating physical reasoning and dynamics accuracy in generative video world models.
Empirical evaluation of domain-adapted language models versus general-purpose LLMs for cybersecurity threat modeling tasks.
Policy gradient method analysis for reinforcement learning in non-Markovian environments with internal state management.
Algebraic framework for latent action models in vision-language-action robotics using action-free video data.
Sparse latent steering technique for interpretable control of molecular editing properties in LLMs with explicit property handles.
Multi-phase pretraining framework for learning representations from electronic health records using joint-embedding predictive architecture.
Training-free method for aligning LLM behavior across cultures without fine-tuning or model access, using black-box prompting techniques.
Pi-Serini search agent system pairing lexical retrieval with frontier LLMs for agentic research tasks, evaluating if BM25 suffices for modern reasoning-capable agents.
Multimodal benchmark dataset with 18,000 samples for evaluating AI-assisted CAD program generation from images and 3D observations.
Benchmark for evaluating LLMs and agents on virtual cell modeling tasks, testing in silico phenotypic screening and biological discovery prediction.
Formal framework for shielding techniques in probabilistic Markov decision processes, extending safety guarantees for autonomous agents with acceptable failure probabilities.
Analysis of on-policy distillation for training reasoning models, investigating when teacher-student supervision helps or hurts performance on token-level tasks.
Study of autonomous data engineering for ML systems, automating dataset discovery, adaptation, and validation to reduce manual data engineering workflows.
Research on engineering robustness into AI agents by applying traditional software engineering processes like testing, adversarial evaluation, and staged deployment instead of on-the-fly synthesis.
Continuous diffusion language models using minimal adaptation to match effectiveness of leading discrete-token language model approaches.
PolyMATH benchmark with 5,000 images evaluating multimodal LLM visual comprehension and abstract reasoning across 10 cognitive challenge categories.
DSGBench evaluation platform for LLM-based agents in strategic games assessing long-horizon reasoning, multi-agent interaction, and decision-making.
Framework using LLMs to automate energy-aware refactoring of parallel scientific code focusing on energy efficiency beyond execution time.
LLM-augmented retrosynthesis system for chemical synthesis and drug development combining ML and LLMs to navigate combinatorial pathway space.
Scalable Bayesian planner for multimodal theory-of-mind reasoning inferring beliefs and intentions without task-specific priors.
Planning-based framework for efficient LLM collaboration combining large and small models to reduce inference costs while maintaining performance.
HAMLET framework combines hierarchical multi-agent LLMs for interactive theatrical experiences with embodied interaction and initiative.
Study of algorithmic recourse for ML-driven decisions addressing multi-stakeholder scenarios with shared constraints.
Systematic comparison of reasoning vs non-reasoning LLMs in judge role evaluating accuracy, efficiency, and robustness on small models.
Empirical scaling laws for language model merging showing power law relationship between model size, expert number, and merging performance.
CritPt benchmark evaluates LLM reasoning on complex open-ended frontier physics research challenges beyond high-school math and coding.
Framework using situational judgment tests and multidimensional item response theory to measure stable behavioral tendencies in persona-conditioned LLMs.
Consensus sampling algorithm aggregates multiple probability distributions to improve generative AI safety with architecture-agnostic approach.
Analysis of neural complex query answering over knowledge graphs comparing learned patterns with training-free query relaxation strategies.
Benchmark for evaluating outcome-driven constraint violations in autonomous AI agents, addressing safety and alignment in high-stakes deployment scenarios.
Recursive Language Models enable LLMs to process arbitrarily long prompts through inference-time scaling via recursive self-calling over prompt snippets.
Batch-of-Thought method processes related queries jointly to improve LLM reasoning by identifying high-quality reasoning templates and detecting errors through consistency analysis.
Framework for characterizing and measuring homogenization and mode collapse in LLMs as an AI safety concern affecting diversity.
Study on inter-rater reliability limitations of human feedback for mental health LLM evaluation, questioning assumptions in RLHF approaches.
Method using concise geometric descriptions to improve multimodal LLM performance on plane geometry problem solving tasks.
Proposes Controllable Information Production (CIP) as a principled intrinsic motivation framework for training autonomous agents without external rewards.
Research benchmark (SayNext-Bench) examining why LLMs struggle with next-utterance anticipation in dialogue compared to human multimodal understanding.