ProgressVLA: Progress-Guided Diffusion Policy for Vision-Language Robotic Manipulation
Vision-language-action model for robotic manipulation with progress estimation to handle long-horizon cascaded tasks.
Vision-language-action model for robotic manipulation with progress estimation to handle long-horizon cascaded tasks.
Multimodal pretraining approach aligning language and vision using GRPO for improved understanding and generation tasks.
Training-free framework for few-shot medical image segmentation using retrieval and adaptation without semantic correspondences.
Deep learning approach for detecting smart contract vulnerabilities using contrastive learning and granular-ball training with limited labeled data.
NITR framework evaluates whether AI coding agents' repository edits preserve software maintainability beyond behavioral correctness.
Study analyzing risks and efficacy of AI-powered facial unmasking for biometric identification, showing it unsuitable for reliable identification.
Survey of counterfactual explanation algorithms for time series classification covering instance-based, pattern-driven, and gradient-based approaches.
CAIAMAR agentic framework for context-aware image anonymization of PII in street-level imagery using multi-agent reasoning.
KVSculpt applies distillation-based approach to KV cache compression for efficient long-context LLM inference combining quantization and merging.
ImagenWorld benchmark with 3.6K condition sets and explainable human evaluation across six image generation and editing tasks.
Luce Alignment Model uses revealed preference techniques to study whether AI agents implement human preferences or pursue independent objectives.
Mat3ra-2D open-source framework for AI-ready design of realistic 2D materials, interfaces, and defects for materials science ML models.
ITQ3_S presents 3-bit weight quantization for LLMs using interleaved ternary quantization with rotation-domain smoothing for efficient inference.
Comprehensive survey of adversarial attacks on multimodal large language models integrating text, images, audio, and video modalities.
Physics-Guided Transformer introduces physics-aware attention mechanism for physics-informed neural networks solving PDEs from sparse observations.
Benchmark for Japanese scene text understanding in vision-language models addressing mixed scripts, vertical writing, and large character inventory.
Benchmark evaluating commonsense-driven hallucinations in vision-language models when visual evidence conflicts with commonsense knowledge.
Federated learning approach using flow-matching generation for privacy protection and robust aggregation against adversarial attacks.
Diffusion-based method for dataset distillation achieving lossless data concentration for efficient training and privacy preservation.
LLM-based agent system for generating interactive documents through human-agent collaboration with improved output control mechanisms.
Analysis of prompt injection attack stages across LLM agent defense mechanisms using cryptographic canary tokens to track attack kill-chain progression.
Unified simulation infrastructure for air-ground embodied AI combining drone and vehicle dynamics in coherent CARLA environment.
Vision-language model pointing mechanism using visual token selection instead of coordinate generation for improved grounding capability.
Pipeline using Vision-Language Models for transcription and semantic annotation of scanned Italian parliamentary speeches.
Hybrid quantum-classical framework combining pretrained EEG encoder with differentiable quantum architecture search for brain signal classification.
Analysis of cultural perspectives embedded in Claude's Constitutional AI through evaluation on World Values Survey items.
First multilingual benchmark for document parsing across digital and photographed documents in diverse scripts and low-resource languages.
LoRA-based domain adaptation method for semantic segmentation using Vision Foundation Models with QR-decomposition.
Benchmark evaluating privilege usage and security risks of LLM agents equipped with real-world tool access across attack surfaces.
ERPO: token-level entropy-regulated policy optimization for reinforcement learning in large language models with improved credit assignment.
DiffAttn: diffusion-based framework with LLM semantic reasoning for drivers' visual attention prediction in autonomous vehicles.
MR-ImagenTime: multi-resolution time series forecasting via hierarchical decomposition and multi-scale conditional diffusion process.
Analysis of categorical perception in LLM hidden states: geometric warping at digit-count boundaries across six model architectures.
Model merging approach for adapting multilingual LLMs to low-resource languages by adding target language weights without retraining.
Pre-deployment framework for estimating learning complexity and communication costs in federated perception systems before training.
Self++: design blueprint for human-AI symbiosis in extended reality preserving human agency while leveraging AI agent capabilities.
EvidenceNet: LLM-assisted framework building disease-specific knowledge graphs from biomedical literature with structured evidence nodes.
Using multimodal LLMs to improve amodal completion for reconstructing occluded regions in images for autonomous vehicles and robotics.
Program analysis framework crossing NL/PL boundary for tracking data flow through LLM API calls in code, enabling taint and slicing analysis.
Analysis of LLM reasoning failures: Bidirectional Coherence Paradox showing competence and grounding dissociate in contemporary LLMs.
Marco DeepResearch: verification-centric design for deep research agents conducting autonomous multi-step reasoning across diverse sources.
First systematic evaluation of membership inference attacks against large audio language models, examining train/test distribution shifts.
Deep reinforcement learning approach for maritime coverage path planning on irregular hexagonal grids without critic networks.
Survey on network performance modeling using simulation and deep learning approaches for predicting packet flow traffic in networks.
EdgeDiT hardware-efficient diffusion transformers optimized for mobile NPUs enabling on-device image generation with reduced compute and memory.
Evolutionary framework using LLMs as generative operators to discover novel reinforcement learning algorithms by searching over executable update rules.
KGroups univariate max-relevance min-redundancy feature selection algorithm for high-dimensional biological data analysis.
AceleradorSNN neuromorphic system integrating spiking neural networks with FPGA for low-latency energy-efficient object detection in autonomous systems.
Federated incremental learning framework for healthcare with dynamic memory replay allocation handling non-IID data across distributed agents.
HISA hierarchical indexing method for fine-grained sparse attention reducing O(L²) bottleneck of token-level sparse mechanisms like DeepSeek Sparse Attention.