The Internet of Physical AI Agents: Interoperability, Longevity, and the Cost of Getting It Wrong
arXiv: Framework for Internet of Physical AI Agents addressing interoperability, security, and sustainability in IoT environments.
arXiv: Framework for Internet of Physical AI Agents addressing interoperability, security, and sustainability in IoT environments.
arXiv: Privacy-preserving federated learning for Alzheimer's classification using 3D MRI with site-aware techniques.
arXiv: Practical guide to AI-assisted research in mathematics and ML, covering productive tool use and responsible guardrails.
arXiv: Analysis of 10,469 experiments by Claude Opus and Gemini agents across 108k design space cells for ML architecture search.
arXiv: VIBEPASS empirically evaluates LLM self-diagnosis and repair capabilities for autonomous software engineering.
arXiv: Benchmarking causal discovery algorithms on synthetic healthcare data for fairness and utility evaluation.
arXiv: LLM-guided neural architecture search for multimodal time-series classification under data-locality constraints for healthcare.
arXiv: LLM family with dynamic tokenizers eliminating fixed vocabulary constraints, up to 70B parameters, improved domain/language adaptation.
MobileLLM-Flash methodology designs on-device LLMs optimized for latency constraints using hardware-in-the-loop architecture search.
ExpertGen automates expert policy generation in simulation for scalable sim-to-real robotic behavior cloning transfer.
MoLoRA enables per-token adapter routing for multimodal generation and mixed-capability requests in multi-adapter serving.
Lightweight proxy models reduce LLM query costs and latency 100x for AI-augmented SQL operations.
Physics-based preprocessing framework standardizes heterogeneous medical images at scale for improved model generalization.
Multi-task RL with chain-of-thought prompting aligns paralinguistic understanding and generation in speech LLMs.
RadAnnotate uses LLMs with retrieval augmentation and selective automation for efficient radiology report annotation.
FormulaCode benchmark evaluates LLM coding agents on repository-level codebase optimization with realistic multi-objective constraints.
Probing-based analysis of moral reasoning trajectories in LLMs across six models showing systematic multi-framework deliberation.
Critic-free RL approach for cross-user activity recognition from wearable sensors with temporal feature generation.
Framework adapts vision-language models as online reward generators for robotic reinforcement learning policy refinement.
Survey of resource consumption threats in LLMs including excessive generation, covering efficiency challenges for providers and users.
HEAR framework extends vision-language-action models to incorporate real-time sound for robotic manipulation tasks.
RecBundle proposes geometric framework for recommender systems addressing information cocoons through topological representation learning.
Inference-time repair layer for retrieval-grounded QA using answer-conditioned counterevidence retrieval to fix commitment errors.
Parallel in-context learning method reducing latency in vision-language models by decoupling demonstration processing from query encoding.
LLM serving system optimizing agentic workflows by handling cross-call dependencies and redundancy from speculative execution.
Data curation method for calibration in LLM compression via frequency-based selection for pruning and quantization.
Local-first multi-agent architecture for automated repository code review with LangGraph orchestration and structured analysis.
Automated skill distillation and adaptation method for financial reasoning in LLMs without fine-tuning.
Reference-free evaluation framework for pathology vision-language models to detect hallucinations without ground truth.
Benchmark for repository-level code understanding with executable environments, enabling agentic code automation tasks.
Benchmark comparing generative augmentation strategies (GANs, diffusion) for bias correction in imbalanced classification under low-data conditions.
Constrained RL method for enforcing hierarchical instruction priority in LLMs via system prompt compliance.
Transformer architecture for 4D point cloud video understanding with temporal scale invariance.
RL method preserving diversity in LLM reasoning via dynamic Jensen-Shannon replay to improve sample efficiency and avoid mode collapse.
Open-source reproduction of Corrective RAG replacing proprietary components with Wikipedia API and open models for improved reproducibility.
Local-first long-term memory system for AI assistants with vector and keyword retrieval, implemented in Rust for conversational agents.
Benchmark and method for evaluating 360° image perception in multimodal LLMs, addressing geometric distortion and spatial reasoning challenges.
Domain adversarial training approach for robust AI-generated audio quality assessment without spurious correlations.
Scoping review of AI-driven digital mental health interventions including GenAI and HCAI across screening, support, and monitoring.
CoMAI multi-agent framework with task decomposition for robust and fair interview evaluation using coordinated LLM agents.
Technical review and taxonomy of 13 generative systems for quantum circuit and quantum code generation including agentic approaches.
Visual prompt discovery method to diagnose and mitigate LVLM perception failures through semantic exploration.
Vision-language process reward models with explicit visual premise verification for reliable step scoring in reasoning.
Genetic programming with surrogate models for dynamic multi-mode project scheduling with simulation-based optimization.
VisBrowse-Bench benchmark for evaluating visual-native search in multimodal browsing agents using MLLMs.
End-to-end framework using Speech LLMs for spoken question answering with attention-guided evidence grounding.
Human-centered architecture for integrating LLM-based cognitive assistants into manufacturing quality management systems.
Security research on sentiment steering attacks targeting RAG-enabled large language models and LLM robustness.
Machine learning pipelines for radio astronomy data processing with explainability focus on automating configuration.
YOLO-based deep learning for automated wasp identification with explainable AI integration for taxonomic classification.