ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?
ViGoR-Bench evaluates visual generative models on reasoning tasks requiring physical, causal, and spatial understanding beyond generation fidelity.
ViGoR-Bench evaluates visual generative models on reasoning tasks requiring physical, causal, and spatial understanding beyond generation fidelity.
Theoretical analysis of simplicity bias in neural networks using Minimum Description Length principle and compression-based learning framework.
GazeQwen integrates gaze information into multimodal LLMs for video understanding via parameter-efficient hidden-state modulation.
Analysis of activation-based safety probes for detecting AI misalignment, proving polynomial-time probes fail on coherent misalignment cases.
Survey of knowledge graph construction methods from unstructured text across domains including healthcare, news, and social media.
GUIDE benchmark for evaluating GUI agents on open-ended tasks emphasizing user collaboration and iterative refinement over automation.
Design principles for integrating resilience and human oversight into LLM-assisted digital twin modeling workflows.
Multimodal Coherence Score metric evaluates fusion quality in multimodal AI systems independent of downstream task accuracy.
Method reinforces structured chain-of-thought in multimodal LLMs for video understanding using reinforcement learning without costly CoT annotation.
Comparative evaluation of sub-10B parameter models on legal reasoning benchmarks using five prompting strategies, testing practical alternatives to frontier LLMs.
Vision-language learning approach for autonomous driving using multimodal infraction datasets to reduce collision rates in end-to-end models.
Study evaluates MedGemma LLM robustness on medical QA benchmarks, finding chain-of-thought prompting reduces accuracy by 5.7% versus direct answering.
PiJEPA framework combining learned navigation policies with world model planning for language-conditioned visual navigation in embodied AI.
Parameter-efficient fine-tuning method for multimodal LLMs addressing fairness disparities across demographic groups in clinical settings.
Benchmark evaluating large vision-language models on zero-shot human age estimation from facial images without task-specific training.
Mechanistic framework to identify and defend against hallucinations in LLMs by analyzing hidden-state dimensions of transformer models.
Behavior-based test revealing selective deficits in LLM theory of mind capabilities, specifically mental self-modeling.
Method for learning variable-sized, data-driven patches via reinforcement learning for compact representation of long-horizon sequence data.
Diary study with 10 knowledge workers examining dependency on LLMs and workflow disruption when LLM access is unavailable.
Multi-agent LLM system for dermatological diagnosis offering interpretability and traceability for clinical reasoning on rare skin diseases.
Analysis of how self-supervised Vision Transformers discover objects, addressing spurious activations in [CLS] token attention for improved object localization.
SWE-PRBench: Benchmark of 350 pull requests showing frontier LLMs detect only 15-31% of human-flagged code review issues, demonstrating gap between code generation and code review performance.
Time-consistent benchmark methodology for evaluating repository-level software engineering systems, addressing synthetic task design and temporal contamination issues.
Philosophical exploration of whether LLMs' sparse auto-encoders suggest a meta-semantic picture of how language models capture meaning.
Evaluation of discrete diffusion vision-language models for GUI grounding in agentic systems, comparing against autoregressive VLMs.
Study of security challenges in open agentic systems combining LLMs with external capabilities, persistent memory, and privileged execution used in coding assistants and enterprise automation.
LLM-based prompting framework automating domain-driven design activities including ubiquitous language, event storming, and bounded contexts.
Method for incorporating conversational context into LLM-based speech recognition through abstract compression of multi-turn audio.
Framework for disability-centered human-agent collaboration proposing channelling, coordinating, and co-creating interaction layers.
Mixed-resolution vision transformer using adaptive token allocation for efficient dense feature extraction from images.
Analysis of late interaction retrieval models, investigating length bias and similarity distribution behaviors in multi-vector scoring.
AI agent system for detecting smart contract vulnerabilities by leveraging DeFi semantics and auditing knowledge summarization.
Physics-aware conditioning scheme for generative video models to enforce physical constraints during generation.
Technique for merging multiple LoRA modules while preserving task-critical directions and addressing subspace misalignment in multi-task adaptation.
Method for merging LoRA-adapted model checkpoints without joint training, handling heterogeneous tasks like classification and regression.
Mechanistic study of how LLMs internally represent and use spatial information, examining whether spatial reasoning relies on structured representations or linguistic heuristics.
Input-adaptive depth aggregation method to mitigate reasoning performance degradation in vision-language model fine-tuning.
CALRK-Bench: context-aware legal reasoning benchmark for Korean law evaluating complex norm interactions.
Information-gain-driven verification approach to reduce hallucinations in multimodal LLM long-form generation.
Generative score inference method for uncertainty quantification on multimodal supervised learning tasks.
AI-driven quantum algorithm discovery using LLMs and evolutionary processes for molecular ground state problems.
Survey of generative modeling in protein design covering neural representations, conditional generation, and evaluation standards.
Analysis of chain-of-thought faithfulness in open-weight reasoning models comparing thinking tokens vs final answers.
Conformal prediction method under covariate shift using kernel mean matching for uncertainty quantification.
CPUBone: vision backbone architecture design optimized for CPU devices with low parallelization capabilities.
ManagerWorker: two-agent pipeline study where expensive model directs cheap model for software engineering tasks.
Neuro-symbolic approach to process anomaly detection combining neural networks with domain knowledge from event logs.
Boltzmann-machine-enhanced Transformer for DNA sequence classification with latent structure discovery.
UNIFERENCE: discrete-event simulation framework for developing and benchmarking distributed AI inference algorithms.
Modality-aware scheduling system for efficient multimodal LLM inference handling text, images, and video workloads.