Causality Is Key to Understand and Balance Multiple Goals in Trustworthy ML and Foundation Models
Framework integrating causal methods into ML to balance trustworthiness objectives: fairness, privacy, robustness, accuracy, explainability.
Framework integrating causal methods into ML to balance trustworthiness objectives: fairness, privacy, robustness, accuracy, explainability.
Re2 dataset for peer review and rebuttal discussions to support automated manuscript evaluation and improve scientific review systems.
AdaBoN prompt-adaptive strategy for Best-of-N alignment reducing computational cost by accounting for alignment difficulty variance.
Guided Policy Optimization framework co-training guider and learner for reinforcement learning in partially observable environments.
Survey on integrating TinyML and LargeML for 6G networks enabling smart healthcare, autonomous vehicles, and digital twins.
Analysis of gender entropy bias in popular LLMs with new benchmark dataset RealWorldQuestioning across business and health domains.
DriveMind framework integrating dual vision-language models with reinforcement learning for interpretable autonomous driving with semantic rewards.
AI Search Paradigm introducing modular architecture of four LLM-powered agents for adaptive information retrieval and multi-stage reasoning tasks.
Improvements to residual reinforcement learning with uncertainty estimation for sample-efficient policy adaptation with sparse rewards.
Fine-tuning approach to align LLM agents with rational and moral preferences in strategic environments, addressing deviations from payoff-sensitive behavior.
Research on LLM emotional reasoning capabilities, showing fragile cognitive understanding of human emotions beyond surface-level recognition tasks.
Object-centric RL approach using dynamic object tokens for visual generalization without reconstruction or auxiliary losses.
Medical continual learning method using unified prompt pools to enable sustainable adaptation while mitigating domain bias and catastrophic forgetting.
Method for achieving robust fine-tuning of non-robust pretrained models using epsilon-scheduling to mitigate suboptimal transfer.
DataMind framework for training scalable data-analytic agents using open-source models with synthetic data generation for multi-step reasoning.
Layer-wise analysis method to distinguish recall from reasoning mechanisms in transformer models via attention and activation patterns.
Safety-constrained reinforcement learning approach integrating Control Barrier Functions during training to enforce dynamic safety constraints.
Method for identifying token-level points for efficient LLM ensembling by aggregating probability distributions for stable long-form generation.
Mathematical proof that transformer language models are injective, enabling exact input recovery from continuous representations.
Benchmark framework for evaluating neural compression and embedding learning methods in Earth observation data.
Method for 3D spatial reasoning from limited views using vision-language models with geometric understanding.
Tutorial on cognitive biases in agentic AI systems for 6G autonomous networks emphasizing true autonomy beyond KPI optimization.
Benchmark framework for evaluating robot policies via real-to-sim translation enabling scalable and reproducible testing of diverse tasks.
Unified image restoration framework using frozen MLLMs with frequency-aware planning for handling multiple degradation types.
Study of catastrophic forgetting in multimodal LLMs under scenario shifts with new MSVQA dataset covering diverse visual conditions.
Method for improving LLM consistency and reliability in business-critical domains using Group Relative Policy Optimization to handle semantic equivalence issues.
Research on model collapse mitigation through AI ecosystem diversity, showing increased diversity across language models prevents knowledge degradation.
Training-free framework decoupling medical knowledge content from delivery for contextualized LLM responses.
Multimodal medical LLM for chest X-ray interpretation with anatomy-aware grounding and spatial reasoning.
Audit study of graduate CS student preferences for AI collaboration in academic tasks.
Methods for automating ontological knowledge base development using LLMs for domain-specific knowledge structuring.
Unified visual encoder supporting both image understanding and generation tasks with VAE-ViT architecture.
Benchmark evaluating vulnerabilities of LLM-based web agents when processing malicious URLs.
Topological state-space networks for higher-order graph learning on combinatorial complexes.
Mixture-of-heterogeneous-experts approach for multivariate long-term time series forecasting.
Research integrating Koopman operators with transformers for time series forecasting with spectral control.
Research on LLM-driven recommendation systems using psychological motivation factors beyond interaction signals.
Research on latent reasoning for chemical LLMs, replacing explicit CoT with continuous structural representations.
VideoTemp-o3 presents an agentic framework for long-video understanding using temporal grounding and thinking-with-videos paradigms. Addresses hallucinations in video analysis through localize-clip-answer pipeline with adaptive frame sampling.
Automated domain-adaptive query expansion framework using LLMs with in-domain exemplar construction and clustering-based selection.
Benchmarking framework using Roofline analysis for characterizing small language models on resource-constrained edge hardware.
Semantic caching system for tiered LLM architectures using asynchronous verification to reduce inference cost and latency.
Probabilistic framework for cost-optimized LLM inference through cascading and routing of models with varying capabilities.
Tool-augmented video summarization system for long-form movie synopsis generation using vision-language models.
Multi-agent AI framework for adaptive AR-based industrial robot training that responds to diverse learner cognitive profiles.
Semantic-visual fusion framework improving multimodal LLM perception of fine-grained visual details through multi-scale context.
Framework combining foundation models and imitation learning for open-vocabulary robot skill adaptation using natural language.
Attack method exposing vulnerabilities in safety alignment of open-source LLMs through deep attention head manipulation.
Analysis of safety mechanisms in LLMs proposing disentangled recognition and refusal axes to explain jailbreak persistence.
Framework for multi-modal controllable long-range visual consistency in generative video content across extended sequences.