NeuroWeaver: An Autonomous Evolutionary Agent for Exploring the Programmatic Space of EEG Analysis Pipelines
Autonomous agent using evolutionary algorithms to optimize EEG analysis pipelines for resource-constrained clinical environments.
Autonomous agent using evolutionary algorithms to optimize EEG analysis pipelines for resource-constrained clinical environments.
Investigates data leakage vulnerabilities in multi-agent LLM systems lacking proper access control and safety mechanisms.
Studies how LLM-powered web agents handle user resources and potential data leakage when automating tasks across live websites.
REMem: episodic memory system for language agents enabling spatiotemporal context recollection and reasoning over interaction histories.
OpAgent: autonomous web navigation agent handling real-world website complexity via supervised fine-tuning and offline RL with online adaptation.
Study on source credibility: LLMs trust human expert feedback more than feedback from other LLMs, analogous to human social influence patterns.
Differentiable rule induction from raw sequence inputs combining ILP with neural networks for interpretable rule learning without symbolic datasets.
Multi-agent proof sprint combining rapid draft generation with adversarial verification and targeted repair for research-level mathematical problems.
Hippocampus: memory management system for agentic AI using binary signatures for semantic search and token-ID streams for efficient persistent storage.
Analysis of quantization paradox where reducing numerical precision from 16-bit to 8/4-bit increases energy consumption while degrading multi-hop reasoning accuracy.
Entropy-based framework for coordinating heterogeneous multi-agent LLM systems through understanding assessment and experience retrieval.
Framework for autonomous GUI navigation agents using MLLMs with Q-estimation and step-wise policy optimization for non-stationary environments.
HyFunc reduces LLM function call latency for agentic AI via hybrid-model cascade and dynamic templating, eliminating redundant processing of function descriptions.
AllMem: hybrid architecture combining Sliding Window Attention with Test-Time Training memory networks to improve LLM efficiency on long-sequence tasks.
PhGPO: Pheromone-guided policy optimization for long-horizon tool planning in LLM agents addressing combinatorial exploration through reusable information.
Automated lightweight AI pipeline using LLMs to solve research-level mathematical problems through natural-language prompting and auto-formalization.
Foundation model approach for relational databases using in-context learning to predict across heterogeneous tabular data without retraining.
OneLatent framework compressing chain-of-thought reasoning into single latent token via supervision from rendered CoT images and OCR hidden states.
Multi-agent research framework using evolutionary search with structured hypothesis management for automated algorithm discovery in experimental domains.
Framework for coordinating multiple independent foundation models to integrate complementary capabilities without retraining black-box systems.
Vashista sparse attention mechanism for constant-time long-context LLM decoding with exponential guarantees by modeling attention as convex hull projection.
End-to-end agentic pipeline using CrewAI-style agent teams for smart contract generation from natural language with automated security and compilation checks.
Framework for A/B testing using content-aware ranking to prioritize variants and extract interpretable insights from historical results.
Framework for claim-level auditability in research agents enabling verification of which passages support which sentences with conflict detection.
Hierarchical reinforcement learning integrating Hindsight Experience Replay mechanism (MOC-HER) for multi-goal sparse reward environments.
LaySPA: Reinforcement learning framework adding explicit spatial reasoning to LLMs for content-aware graphic layout design as structured policy learning problem.
Hybrid memory architecture with dynamic retrieval scheduling for LLM agents balancing efficiency and effectiveness in extended dialogues without lossy compression.
Statistically principled early stopping methods monitoring uncertainty signals during LLM reasoning generation to prevent unnecessary computation from overthinking.
Framework analyzing streaming external memory for LLMs where insertions and retrievals interleave, covering ingestion, maintenance, retrieval and integration of information.
Method for compressing long contexts into soft prompt embeddings using block-wise causal masking to reduce LLM inference latency from quadratic self-attention complexity.
Methods for aligning AI clinical reasoning with structured frameworks using abductive explanations to improve trust and identify critical symptoms.
Modular agent framework with generative reasoning for deploying large AI models on edge devices with improved flexibility and reduced computational demands.
FloCA framework for flowchart reasoning in dialogue systems ensuring logical consistency and faithful node transitions in multi-turn interactions.
Adaptive memory structures for LLM agents that dynamically select memory organization based on interaction context for improved long-horizon performance.
REAL framework for resolving knowledge conflicts in visual question answering using reasoning-pivot alignment for conflict detection.
Plan-MCTS framework combining tree search with LLMs for improved web navigation by addressing sparse valid paths and noisy context challenges.
GUI-GENESIS framework for automatically synthesizing efficient GUI training environments with verifiable rewards for post-training agent development.
Research on steganographic chain-of-thought in LLMs where models hide reasoning in text, evaluating safety risks across 28 models.
Proposes algebraic quantum intelligence framework to improve creative output in LLMs by addressing structural constraints on generation spaces.
ForesightSafety Bench evaluates frontier AI risks across multiple dimensions addressing systemic risks from increasingly autonomous goal-directed AI systems.
Multi-agent RL system for clinical reasoning emphasizing process-grounded decision-making aligned with medical standards over outcome-only accuracy.
Text-guided staged knowledge injection improves agentic RL with verifiable rewards for ultra-high-resolution remote sensing understanding.
CORPGEN simulates corporate environments with autonomous agents managing 45+ concurrent long-horizon tasks with interleaving and dependencies.
REDSearcher framework optimizes LLMs for deep search tasks by addressing sparse high-quality trajectories and high interaction costs in long-horizon reasoning.
GRAIL framework uses imitation learning to align goal recognition with actual agent behavior for improved AI alignment and intention understanding.
Examines saturation of LLM benchmarks and proposes solutions as frontier models increasingly saturate new evaluations shortly after publication.
Analyzes atomic-level mechanisms of potentially dangerous behaviors in locally-deployed LLMs and proposes detection methods without cloud connectivity.
PITA dataset with 23M propositional logic statements evaluates how reasoning traces support length generalization in neural networks.
Precedent Informed Reasoning reduces computational costs and improves performance in LLMs by leveraging past similar cases to constrain reasoning search spaces.
Comprehensive frontier AI risk assessment framework evaluating risks from advanced LLMs and agentic AI across five critical dimensions.