TRACE: A Two-Channel Robust Attribution Watermark via Complementary Embeddings for LLM-Agent Trajectories
Watermarking scheme for LLM-agent trajectory logs enabling attribution verification against resellers with full data access.
Watermarking scheme for LLM-agent trajectory logs enabling attribution verification against resellers with full data access.
Disease-aware generative language model for drug discovery that conditions molecular generation on disease ontology and protein sequences.
Uses Group Relative Policy Optimization to adapt LLM-based ASR systems trained on synthetic speech for real-world banking domain.
World Action Models trained on egocentric human video for robot manipulation, separating transferable task semantics from human-specific factors.
Q-learning approach for adaptive model retraining in Open RAN networks to handle traffic-induced performance drift.
Research on LLM abstention mechanisms distinguishing between incorrect answers and unanswerable questions using separate confidence axes.
Research paper analyzing interaction-level disparities in AI agent access beyond availability/quality/quantity dimensions.
Multimodal agent with episodic memory for understanding, generation, and editing without context window explosion in long-horizon dialogue.
Audits reliability variation in LLM-as-judge evaluation across model upgrades, showing judge replacements are not interchangeable measurement tools.
DocMaster hierarchical document analysis system preserving structural relationships for LLM-based analysis of academic papers and technical documents.
VocaDet open-vocabulary object detection and segmentation via visual tokenization and vector database retrieval for scalable category expansion.
SMetric session-centric LLM scheduling system optimized for agentic serving with high KV-cache reuse, balancing throughput and latency.
Structured sparse autoencoders for learning modality-consistent concepts across vision-language models with improved mechanistic interpretability.
UltraX adaptive programmatic editing system for large-scale pre-training data refinement beyond rule-based and rule-learning approaches.
Multi-modal machine teaching approach for robust reward learning in autonomous agents across diverse operational environments using inverse RL.
WebSwarm multi-agent orchestration system for deep-and-wide web search using recursive agent coordination beyond single trajectory limitations.
Training-free speculative decoding acceleration for LLM sampling with relaxed distribution guarantees enabling speed-capability trade-offs.
ProjAgent retrieves procedurally similar repository functions for repository-level code generation, accounting for cross-file dependencies and project conventions.
Validity assessment of Portugal's AMALIA 9B LLM for data annotation tasks, comparing agreement and reliability against human coders.
In-training low-rank regularization technique for neural network compression without requiring SVD or architecture modification.
Chain-of-Frame reasoning approach using video generation models as alternative to chain-of-thought for logical reasoning in large models.
Multi-perspective causal discovery framework using LLMs for abductive reasoning, includes DeepAbduction dataset for pollution cause analysis.
Curriculum learning method for Direct Preference Optimization using two-dimensional difficulty space (prompt complexity and distinguishability) to improve LLM alignment.
Research on energy-efficient domain-specific AI models and agents, addressing computational costs of large language models in production.
System for training task-oriented dialogue agents for recruitment using simulator-based data evaluation and selection.
Case study using ChatGPT for rapid scientific prototyping in lunar trajectory estimation competition, achieving second place.
Benchmark for evaluating long-horizon autonomous decision-making and strategy stability of LLM agents in supermarket simulation.
Position paper arguing RAG systems need redesign to handle opinion-rich content beyond factual grounding.
Neural architecture for learning long-range non-stationary temporal patterns in streaming settings without revisiting past data.
Study of memory design in long-lived foundation model agents, analyzing personalization, extraction risk, and deletion fidelity tradeoffs.
Lightweight LLM-based agent framework for rare disease diagnosis built through policy iteration with human feedback.
Tool for formal verification of neural ordinary differential equations in safety-critical applications.
Multi-layer benchmark with 118 problems for evaluating LLMs on reconstructing and applying expert investment decision frameworks.
System prompt technique (narration-of-thought) for improving LLM ethical reasoning on moral dilemmas by reducing stakeholder collapse.
Token-efficient context management module for LLM agents performing repository-level program repair with precision evidence selection.
Analysis of security vulnerabilities in persistent-state AI coding agents that ship code iteratively across sessions.
Benchmark for evaluating LLM agents across multilingual long-horizon tasks requiring planning, tool use, and environment interaction.
Theoretical framework for embedding cognitive architectures natively into LLMs rather than simulating via prompting and context management.
Expert-based analysis of explainable AI methods for safe AI development and certification under regulatory frameworks.
ParamMute method suppresses knowledge-critical feed-forward networks to improve faithfulness in retrieval-augmented generation systems.
Synthetic data generation framework for personalizing vision-language models using concept hierarchies.
Large-model-driven semantic multiple access scheme for token-based communications integrating pretrained models.
Thunder-Tok tokenizer reduces token fertility for language models while maintaining performance and inference efficiency.
Open-source evaluation framework for topological mapping systems with standardized metrics and perceptual aliasing quantification.
Adaptive framework for generating bias-eliciting questions to evaluate and expose biases in large language models.
Federated learning framework reducing memory and communication overhead for edge AI deployment in wireless networks.
Value-guided planning framework for vision-language-action robotic models that improves performance under distribution shift and long-horizon tasks.
Small language models can serve as efficient reference-free evaluators by using internal representations instead of generation.
Study of feature entanglement in language models and methods for isolated interventions via almost orthogonal feature representations.
Self-EvolveRec uses LLM-based directional feedback for automated recommender system design in open-ended program spaces.