ForestPrune: Training-free token compression for video MLLMs using spatial-temporal modeling. Achieves high-ratio compression for video processing.
EVA: Reinforcement learning method for video understanding agents using multimodal LLMs. Adaptive frame sampling and reasoning without manual workflows.
Method for LLMs to return set-valued predictions with coverage guarantees instead of single outputs. Improves answer discovery through repeated sampling.
Graph foundation models tested for zero-shot generalization across different GNN architectures and scales.
Visual backdoor attacks exploit mobile GUI agents via notification-based remote action execution.
Tabular data generation via probabilistic circuits questioned; current benchmarks overstated progress.
Concept-based explainability framework for flood/wildfire detection models in disaster management.
GLA-CLIP enables training-free open-vocabulary semantic segmentation with global-local window alignment.
Kolmogorov-Arnold networks improve YOLOv10 interpretability for object detection in degraded conditions.
RAG fine-tuning evaluation for EDA long-form generation with novel human evaluation metric TriFEX.
DBAutoDoc automates database schema documentation combining statistical analysis with iterative LLM refinement.
AuthorMix uses modular layer-wise adapters for lightweight, flexible authorship style transfer with meaning preservation.
LLMs detect microservice architecture patterns across multiple programming languages outperforming single-language tools.
Explainable AI analysis reveals AI-generated text detectors exploit dataset artifacts rather than genuine detection signals.
Activation watermarking technique detects adaptive adversarial attacks against LLMs attempting to evade safety monitoring.
Semantic ID tokens enable LLM-based generative recommendation systems with efficient decoding over large item corpora.
Implicit reward modeling from human feedback like clicks for cost-effective LLM alignment via RLHF.
Foundational ML theory for learning under regime variation with evolving learner state and evaluation conditions.
Investigation of neural ODEs and SDEs for model-based reinforcement learning, showing neural SDEs better capture stochasticity in environment dynamics.
WeCAN: reinforcement learning framework for heterogeneous DAG scheduling addressing task compatibility, resource constraints, and rapid schedule generation.
SafeSeek framework for universal attribution of safety circuits in LLMs using mechanistic interpretability to understand alignment, jailbreak, and backdoor behaviors.
Query-efficient jailbreak fuzzing method for LLMs that identifies token importance during prompt mutation to reduce redundant searching under query constraints.
Multimodal framework for human-multi-agent interaction integrating perception, embodied expression, and coordinated decision-making in shared physical spaces.
Analysis of LLM-based social network where autonomous AI agents interact through natural language, studying collective dynamics and emergent network fragility.
Comparative study of seven machine learning models for hourly weather forecasting in complex topography, including XGBoost, LSTM, and CNN-LSTM variants.
Agentic AI platform for portfolio investment screening using LLM agents for fundamental analysis and sentiment analysis with deliberation mechanism for buy/sell signals.
Analyzes user perception of Android's Earthquake Alert system using LLMs on social media data from 2025 Türkiye earthquake.
Surveys natural language interfaces to spatial and temporal databases, covering methods and taxonomy for NLIDBs with geospatial data.
Proposes energy-based modeling for discrete graph generation using transport-aligned sampling to improve efficiency and quality.
Extends Priority Inheritance with Backtracking (PIBT) algorithm for multi-agent path finding with multiple dependencies in congested environments.
Proposes SortedRL, a length-aware scheduling method to accelerate RL training for LLMs, reducing rollout bottleneck in long chain-of-thought generation.
Examines how humans attribute errors in multi-agent AI systems under delayed feedback, revealing biases in decision-making across sequential steps.
Studies practical adversarial attack feasibility against ML-based IoT intrusion detection systems, addressing implementation constraints.
Evaluates whether LLM-generated tests reflect genuine program understanding or superficial pattern reproduction, examining behavior under software evolution.
Proposes 3DCity-LLM, a multimodal LLM framework for 3D city-scale perception using coarse-to-fine feature encoding across object, relational, and global contexts.
Introduces benchmark dataset and evaluation framework for code review agents, addressing code quality assurance as AI-generated code scales.
Proposes VTAM, extending video-action models for embodied AI with tactile sensing for contact-rich physical interactions beyond vision-only approaches.
ReqFusion integrates multiple LLM providers (GPT, Claude, Groq) to automate software requirements extraction, classification, and analysis.
Shows LLMs produce unstable outputs on gender inference tasks under minimal context variations, revealing dependence on cultural stereotypes in training data.
Proposes VISOR, a method to reduce inference costs in large vision-language models through dynamic, sparse vision-language interactions without information bottlenecks.
Evaluates Vision Language Models' ability to perform pre-diagnostic sanity checks in medical imaging, identifying gaps between fluent text generation and safe visual understanding.
BSDS system architecture integrating AI agents with data platforms for business-semantic-centric decision-making and workflows.
TRACE framework for self-evolving agent benchmarks that dynamically increase difficulty using test-time exploration and validation.
BIRD-INTERACT benchmark evaluating LLMs on multi-turn text-to-SQL tasks with dynamic interactions and error handling.
BuilderBench benchmark for evaluating AI agents' ability to learn through exploration and interaction beyond training data patterns.
Hybrid Stackelberg game and diffusion-based auction mechanism for task offloading among collaborative AI agents in Internet of Agents.
Domain-specific risk taxonomy and evaluation framework for LLM-based driving assistants addressing safety-critical scenarios.
Entropy-based analysis shows reducing entropy improves tool-use behavior in LLM agents, reducing excessive tool calls and latency.
Classical Chinese jailbreak prompts bypass LLM safety constraints more effectively than English due to obscurity and conciseness.
Evidence-grounded diagnostic reasoning agent using vision-language models for chest X-ray interpretation.