AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
AlpsBench: evaluation benchmark for LLM personalization using real dialogue data covering memorization and preference alignment tasks.
AlpsBench: evaluation benchmark for LLM personalization using real dialogue data covering memorization and preference alignment tasks.
LITTA framework for multimodal document retrieval using query expansion to improve evidence page retrieval from visually rich documents without retraining.
SpatialPoint framework for embodied AI predicting executable 3D points from visual-language input, enabling physical interaction and navigation in environments.
Physics-Informed Mamba generative model for protein backbone design using spectral initialization and flow matching for efficient long-sequence generation.
Theoretical analysis of divergence between expanding LLM context windows (512 to 2M tokens) and contracting human attention capacity from 2017-2026.
Modular LLM-driven framework for automated candidate assessment in recruitment, integrating job descriptions, CVs, and interviews for structured evaluation reports.
Study of carbon footprint in economic research workflows using generative AI as tool, shifting analysis from model-level to workflow-level impact assessment.
Analysis of benchmarking challenges for multi-agent scientific AI systems including contamination risks, ground truth issues, and tool use complications.
Goal-conditioned offline RL agent for learning surgical needle trajectories from endoscopic video, enabling robot-assisted suturing with trajectory prediction.
Theoretical proof that capability safety is exactly representable as propositional Datalog, enabling algorithmic optimization and incremental maintenance.
Schema-based evaluation and routing system for multi-provider LLM gateways enabling fine-grained quality assessment and intelligent request routing across models.
Systematic analysis of how vision-language models extract scene context from single objects, examining robustness implications through behavioral and mechanistic study.
Distilled LLM-driven Sparse Mixture-of-Experts framework for visual recognition combining text guidance with dynamic expert activation for improved generalization.
Semantic segmentation method incorporating ordinal relationships between classes for improved accuracy on medical and odontological images.
Structured Sequential Visual Chain-of-Thought reasoning for multimodal LLMs, enabling selective sequential visual attention instead of static visual tokens.
Vision-language model for sleep staging from polysomnography waveforms, generating clinician-readable explanations grounded in medical scoring criteria.
Language-conditioned visual navigation using world modeling, where embodied agents follow natural language instructions based on visual observations and continuous control.
Framework integrating Sparse Autoencoders with dynamic head pruning in Vision Transformers for interpretable and controllable model compression.
Enhanced datasets and benchmarking framework for autonomous landing systems, addressing object detection dataset limitations through diverse data sources.
Diffusion-based approach for modeling Pareto set evolution in dynamic multiobjective optimization problems without requiring training.
Pipeline for generating synthetic training data from camera trap imagery to enable ML models for automated wildlife health screening (alopecia and body condition detection).
Compact Vision Transformer for potato leaf disease classification using efficient architecture design and explainability features.
Application of vision-language models to assess handwritten Chinese character quality and generate actionable feedback for learners.
Study comparing failure modes of compressed vision-language models on edge devices, proposing taxonomy of error types (Object Blindness, Semantic Drift, Prior Bias) across different model sizes.
Framework for automated semantic annotation of broadcast television using multimodal LLMs, evaluating different pipeline architectures and input configurations for video understanding.
Method to improve in-context demonstration selection for multimodal LLMs by reformulating it as a sequential task rather than relying on k-NN similarity, improving performance on visual regression tasks.
TED proposes a training-free knowledge distillation framework for multimodal models using context-based approaches instead of parameter optimization, enabling deployment in resource-constrained environments.
arXiv research on spatial reasoning gaps in frontier LLMs, testing external imagery module as cognitive prosthetic for mental rotation tasks.
Federated learning approach for pre-training multimodal large language models on distributed privacy-sensitive data without centralized data collection.
CRISP: LLM-based method to jointly rank cited papers within citing papers by analyzing citation context, mitigating positional bias through randomized ranking.
Open-source reproducible framework for deep learning-based brain tumor classification from multi-sequence MRI with emphasis on accessibility and academic development.
Batch-level routing framework for LLM queries that optimizes model assignment under cost, GPU capacity, and concurrency constraints with resource-aware optimization.
Framework to interpret and verify semantic hierarchies in CLIP vision-language model embeddings through agglomerative clustering and hierarchy alignment.
Neural operators with dual-scale architecture for stable long-term fluid dynamics forecasting by addressing local detail blurring and precision issues in PDE modeling.
Unified sparsification framework for cross-modality prediction across graphs, language, and tabular data using L0-gated representations for efficient multimodal learning.
First approach to vibrotactile captioning, generating natural language descriptions from vibration signals for VR and human-computer interaction applications.
GroupRAG framework combining retrieval-augmented generation with cognitively-inspired group-aware reasoning and knowledge-driven problem structuring.
Hybrid document-routed retrieval approach for RAG systems in financial QA using semantic file routing to avoid chunk confusion.
Analysis of throughput optimization strategies for LLM training, including dataloader and memory profiling innovations for reducing computational bottlenecks.
Activation-patching method to expose hidden hallucinations in language models by surfacing errors suppressed in safety circuits.
Statistical regression framework for analyzing how prompt features impact LLM performance, extending XAI methods for language models.
SpatialAnt: Zero-shot robot navigation system using MLLMs with autonomous scene reconstruction and visual anticipation without pre-built scene priors.
Foundation model approach (VAN-AD) for time series anomaly detection using visual masked autoencoders with normalizing flows for cross-dataset generalization.
GISclaw: Open-source LLM-powered agent system for full-stack geospatial analysis supporting vector and raster data with multi-model architecture.
Method detecting and mitigating intrinsic LLM deception by identifying stability asymmetry between hidden reasoning and generated responses.
Vision-and-Language Navigation framework (BTK) integrating multimodal knowledge bases with generative models for improved semantic alignment in navigation tasks.
EZASP: Developer tool simplifying Answer Set Programming usage through LLM-assisted program generation and natural language interfaces.
Analysis of strategic behavior in AI ranking arenas where model producers submit multiple variants to artificially improve rankings through noisy preferences.
Controlled empirical evaluation of how LLM choice, model size, learning approach, and prompting strategy affect political text annotation results.
Evaluation of LLM capabilities for quantum software, architecture, and system design across multiple quantum computing domains and problem types.