A toy framework for single and multi-agent human-AI curiosity ecosystems
Toy framework modeling curiosity as ecosystem with adjustable weights for uncertainty reduction, costs, and delayed returns.
Toy framework modeling curiosity as ecosystem with adjustable weights for uncertainty reduction, costs, and delayed returns.
Information gain-based adaptive tree-structured rollout optimization for multi-turn LLM agents on long-horizon tasks.
TOFFEE system for synthesizing high-quality data agent trajectories at scale for heterogeneous enterprise environments.
Theoretical framework for embedding application-layer cognitive protocols into native LLM meta-architecture via structural mechanisms.
Task decomposition-guided reranking method for improved skill selection in agent systems with large skill libraries.
LLM safety guardrail using intent-driven reasoning to balance efficiency and robustness against complex safety risks.
Post-hoc interpretability module using dictionary learning to decompose autonomous driving model behavior into semantic concepts.
Training-free framework using agentic topology sampling and knowledge graphs for zero-shot IoT forecasting in building sensor networks.
ExplAIner declarative query language for specifying and analyzing model explanation methods uniformly across classification models.
Multi-agent LLM system extracts H. pylori infection evidence from medical reports across heterogeneous structured and free-text fields.
Danus system orchestrates LLM-based mathematical reasoning agents using fact-graph memory for research-level problem solving.
Multi-agent deep reinforcement learning system for battery management optimization in dairy farm renewable energy integration.
Method to predict LLM agent failures early using probe cascades on hidden activations to avoid wasted compute in multi-step tasks.
Large-scale multivariate time series dataset for training foundation models on real-world data rather than synthetic datasets.
FootsiesGym: open-source benchmark environment for two-player zero-sum imperfect-information games with vectorized simulator for efficient training.
FreqDepthKV: frequency-guided cache compression factorizing KV states into shared low-frequency and sparse high-frequency components for long-context inference.
VAORA: reward design for vision-language models to improve physical reasoning and align internal reasoning with actions in interactive tasks.
DepthWeave-KV: token-adaptive KV cache compression across transformer layers using residual factorization for long-context inference.
Examines AI's impact on linguistic and cultural preservation in Indian subcontinent, viewing AI as double-edged for inclusion vs. homogenization.
PORTICO: reference monitor for revocable capabilities in coding agents, compiling task contracts into initial capabilities and closure predicates.
Runtime verification framework for AI agent actions: formalizes authorization, tamper-evidence, and deterministic reconstruction of agent trajectories.
Benchmark comparing KV-cache optimization techniques (quantization, pruning, merging) across models, tasks, and serving stacks for long-context LLM inference.
Analyzes ICLR papers 2017-2025 to identify trajectory-changing methodological contributions using public reviewer scores and decisions.
CCBENCH: benchmark evaluating LLM cultural competency by assessing adaptation to implicitly signaled cultural norms in health queries.
GAIDE framework enabling K-12 teachers to create AI-powered learning tech via vibe coding with LLMs, supporting teachers as designers.
Uses pattern-based knowledge components to automatically recommend programming learning resources, reducing need for expert curation.
CANONIC: system that applies compiler-like governance to LLM-generated content, using formal grammars to admit/reject artifacts into evidence ledgers at scale.
CHARLIE: on-premise multi-agent RAG system for structured evidential reasoning in digital forensics with traceability and compliance requirements.
Multimodal RAG system using post-hoc selective modality escalation to balance cost and utility by adaptively invoking vision-language models.
PORTS: preference-optimized retrieval method for tool selection in LLM agents. Aligns retrievers with tool-calling LLMs through joint training.
Multi-domain scientific code search benchmark with 5,264 curated repositories across scientific computing domains to evaluate code discovery tools.
Offline RL approach to learn control policies for LLM agent execution harnesses. Formalizes harness operation as finite-horizon MDP with frozen LLM executor.
BioSecBench-Refusal: benchmark for evaluating AI agent safety in biology tasks, pairing 61 routine tasks with 46 red-team scenarios for biosecurity assessment.
CanvasAgent: multi-modal agent orchestrating visual tools for complex image creation and editing through synthesis, segmentation, and composition.
KAT-Coder-V2.5: agentic coding model trained autonomously in executable repositories. Introduces AutoBuilder for sandbox reconstruction and end-to-end agentic post-training framework.
Comprehensive measurement study of mobile LLM inference across frameworks (llama.cpp, GENIE) and hardware backends (CPU, GPU, NPU). Introduces PowerBench profiling tool to identify efficiency bottlenecks.
Study of decision protocols in multi-agent LLM conversations for task performance improvement through specialized agent distribution and discussion mechanisms.
Analysis of confidentiality and legal privilege risks in generative AI systems across training, context windows, and RAG-based knowledge databases.
Calibration method for binary classifiers in adversarial environments maintaining consistent false-positive rates across predictions during model retraining.
PatchOptic system for managing shared state in agentic LLM workflows using projected views and verified structured updates instead of grep-like searches.
Analysis of naturally occurring statistical signals in vision datasets that behave like backdoor triggers, studying Imagenet patterns linked to labels.
aiAuthZ security framework for identity-bound authorization in AI agents, evaluating LLM vulnerability to tool-call forgery attacks across 15 models.
Self-Review Reinforcement Learning method for LLMs using cross-episode memory and policy distillation to handle sparse/delayed environmental feedback.
Study showing LLM conformity to peer responses is largely confounded by repeated answers independent of speaker presence in benchmark tasks.
Analysis of LLM yes-no bias in binary judgments, isolating effects of answer order and wording from moral judgment shifts using psychometric methods.
Evaluation study comparing prompt robustness between objective and subjective LLM tasks, showing different sensitivity patterns across model families.
Training-free approach using generative image models for 3D primitive shape abstraction without task-specific fine-tuning.
Analysis of structural concentration in AI bias research community, examining whose fairness definitions and debiasing frameworks are produced.
ResonatorLM architecture for efficient long-context language modeling using causal resonant field mixing as alternative to transformers.
Continual learning research challenging retention-centered assumptions and prioritizing real-time adaptation in non-stationary environments.