HTM-EAR: Importance-Preserving Tiered Memory with Hybrid Routing under Saturation
HTM-EAR hierarchical memory system for long-running agents combining importance-aware eviction with HNSW-based working memory and archival storage.
HTM-EAR hierarchical memory system for long-running agents combining importance-aware eviction with HNSW-based working memory and archival storage.
Benchmark evaluating graph foundation models across topic and format domain shifts to assess knowledge transfer capabilities.
Research paper introducing Flip-Agent, first framework for targeted bit-flip attacks on multi-stage LLM-based agent pipelines with external tools.
Large-scale controlled study (N=62,808) on how agentic scaffolds affect measured safety in LLM deployments across six frontier models and four configurations.
Research paper on continual learning for wearable sensor activity recognition, addressing catastrophic forgetting in IoT human activity recognition systems.
Research paper comparing five prompt engineering strategies to reduce LLM hallucinations and increase output consistency in industrial high-stakes applications.
Research paper on improved implementation of Sharpness-Aware Minimization (SAM), an optimization technique that enhances model generalization by minimizing loss in parameter neighborhoods.
Python tool implementing Combinatorial Fusion Analysis for ensemble learning classifier generation using rank-score characteristics and cognitive diversity.
Proposes neural cellular automata for synthetic data generation in LLM pre-training to address natural language limitations and reduce human bias.
Agentic AIBOMs extends Software Bills of Materials to capture runtime behavior and reproducibility in AI systems, addressing supply-chain security for dynamic execution.
NabaOS framework for detecting hallucinations in tool-using AI agents via lightweight verification receipts, avoiding expensive zero-knowledge proof overhead.
Position paper addressing memory architecture challenges in multi-agent LLM systems, proposing three-layer hierarchy and identifying protocol gaps for cache sharing and memory access control.
HTMuon optimizer improves upon Muon by preserving heavy-tailed weight spectra for more effective LLM training using spectral correction.
ADVERSA framework measuring guardrail degradation in multi-turn LLM interactions using automated red-teaming with 70B attacker model.
Sparse autoencoders applied to time series foundation model (Chronos) revealing causal feature hierarchies through ablation experiments.
Failure analysis of LLM-generated security patches across 319 examples showing only 24.8% achieve full correctness on security vulnerabilities.
Adversarial semantic layer activation steering technique for red-teaming and suppressing harmful content generation in LLMs.
Multi-agent framework using LLMs to optimize GPU kernels with explicit, interpretable optimization strategies replacing opaque heuristics.
ES-dLLM improves diffusion language model inference efficiency through early-skipping of redundant context processing iterations.
Multi-Stream Perturbation Attack exploits vulnerabilities in LLM thinking mode by concurrent task interference to bypass safety alignment.
Analysis of safety risks in agentic LLM systems using local executors and tool use, focusing on execution-layer attack surfaces and capability supply chains.
Equivariant asynchronous diffusion model with adaptive denoising schedule for 3D molecular conformation generation capturing hierarchical structure.
Code-Space Response Oracles (CSRO) generates interpretable multi-agent policies using LLMs instead of black-box neural networks for game-theoretic equilibria.
CLIPO framework improving LLM reasoning by using contrastive learning in policy optimization to evaluate intermediate reasoning step correctness.
Analysis showing 'Lost in the Middle' performance degradation in LLMs is geometrically inherent at transformer initialization, predating training.
Autoregressive Action Expert for Vision-Language-Action models maintaining long-lived memory context for robot/agent control tasks.
Method for faster LLM finetuning by reusing and remixing existing model checkpoints instead of training from scratch on each new task.
Identifies and exploits MCP clause-compliance vulnerabilities in Model Context Protocol standard for agent-tool integration.
Develops risk assessment framework for open-source MCP servers enabling LLM agents to access external tools securely.
Presents Adaptive Activation Cancellation framework treating hallucinations as interference in transformer residual streams for real-time mitigation.
Proposes Delta-K, inference framework addressing concept omission in multi-instance diffusion image generation.
Analyzes policy gradient for k-armed stochastic bandits using continuous-time diffusion approximation with regret bounds.
Studies harmonic loss with non-Euclidean distance layers as alternative to cross-entropy for neural network training.
Presents DUCTILE, agentic LLM orchestration system for automating engineering analysis in product development workflows.
Designs conversational AI system for querying 1.7M digitized natural history museum specimens using semantic search.
Introduces Simulation-in-the-Reasoning framework embedding domain simulators into LLM reasoning loops for autonomous transportation.
Develops automated benchmark for evaluating novelty of research ideas using LLMs to address literature review scalability.
Vision-Language-Action model improvement via concept-gated visual distillation for robotic manipulation in cluttered scenes.
Federated active learning study addressing class imbalance and non-IID data in privacy-preserving annotation.
Research on translationese bias in multilingual LLM evaluators; proposes information bottleneck solution for low-resource languages.
LLM-based congestion control protocol using utility functions for network applications to optimize sending rates in distributed settings.
Dynamic knowledge fusion approach for multi-domain dialogue state tracking in task-oriented dialogue systems with limited training data.
Attention reformulation for generative recommender systems addressing structural inefficiencies of interleaving tokens in sequence generation.
Few-shot adaptation framework for robots in non-stationary environments using latent trend embeddings to handle concept shift.
Mixed-methods study examining how AI co-writing tools change user engagement with ideas and opinion formation through behavioral analysis.
Causal Concept Graphs method using sparse autoencoders and differentiable structure learning to interpret multi-step reasoning in LLMs.
Neural scaling laws for Mixture-of-Experts models defining optimal compute allocation between expert and attention sub-layers.
Safety framework for human-robot interaction combining control barrier functions with conformal risk control for formal guarantees.
Theoretical analysis of SGD learning dynamics in two-layer linear networks trained with label noise, studying implicit bias mechanisms.
Framework using LLMs to design service systems by analyzing textual evidence from customer support and compliance data to optimize configurations.