SPEAR: An Engineering Case Study of Multi-Agent Coordination for Smart Contract Auditing
Multi-agent framework for smart contract auditing using specialized agents for planning, execution, and recovery with coordination protocols.
Multi-agent framework for smart contract auditing using specialized agents for planning, execution, and recovery with coordination protocols.
Study demonstrating LLM biases when simulating misinformation susceptibility, showing models overstate attitudes and ignore network effects present in humans.
Qualitative study of 33 K12 teachers' perspectives on using conversational AI agents to scaffold group collaboration in classrooms.
Knowledge distillation method for distilling RL-trained LLMs with chain-of-thought reasoning into smaller student models while preserving reasoning capabilities.
Analysis showing chain-of-thought prompting underperforms direct answering in medical vision-language models due to perception bottlenecks in domain-specific tasks.
Parallel framework combining imitation and reinforcement learning for autonomous driving, addressing limitations of sequential fine-tuning approaches.
Method to improve pretrained generative robot policies by replacing sampled noise with optimized constant noise vectors for downstream reward optimization.
Mid-training adaptation strategy for LLMs to improve automatic summarization of radiology reports, exploring domain-specific pre-training approaches.
CAIAMAR: multi-agent framework for context-aware image anonymization in street-level imagery using agentic reasoning.
Kill-chain canary methodology for tracking prompt injection attacks across multi-agent LLM systems with stage-level diagnostics.
System for making mathematical theorems interactive by grounding LLM-generated explanations in formal representations enabling execution and stepping.
Framework for eliciting and verbalizing LLM assumptions to explain and mitigate sycophantic behavior in model outputs.
Multi-stage LLM-assisted workflow for scientific algorithm development separating theory extraction, formal specification, and code generation.
Method for LLM personalization using a small portfolio of models capturing diverse user preferences without per-user models.
Distributional reinforcement learning approach for decision-making in healthcare, accounting for uncertainty across heterogeneous populations.
ALTO: system for adaptive LoRA hyperparameter tuning and orchestration across heterogeneous LLM fine-tuning workloads in multi-tenant environments.
WisdomInterrogatory (LuWen): open-source Chinese legal language model built on Baichuan foundation model for legal domain applications.
System for safe capability evolution in embodied agents with compatibility checking and runtime rollback mechanisms.
Training-free open-vocabulary semantic segmentation framework (OV-Stitcher) leveraging pretrained vision-language models without additional training.
HyperMem: hypergraph-based memory architecture for conversational agents enabling long-term context tracking and high-order associations.
Physics-aligned simulator (SIM1) for generating synthetic data in deformable object robotic manipulation tasks.
Framework combining LLMs with graph neural networks for text-attributed graph learning in low-resource settings using GNN feedback.
QuanBench+ unified benchmark for LLM quantum code generation across Qiskit, PennyLane, Cirq with 42 aligned executable tasks.
Benchmark evaluating robustness of LLM reasoning with 14 perturbation techniques applied to mathematical reasoning tasks.
DRTO combines token-level RLHF with distributional robustness to improve LLM resilience to input perturbations and formatting changes.
Automated label function generation for data annotation using LLMs with structured exploration-exploitation strategy.
TinyML Z-score anomaly detection system running on resource-constrained microcontrollers using power side-channel data.
CSAttention: sparse attention mechanism for accelerating LLM inference by reducing KV-cache bottlenecks through centroid-scoring without retraining.
Framework evaluating when LLMs should act versus escalate decisions using uncertainty estimation across five real-world domains.
AlphaLab autonomous research system using frontier LLMs as agents to automate full experimental cycles in optimization domains without human intervention.
StructRL: recovers dynamic programming structure from distributional RL learning dynamics. Bridges data-driven and structured approaches for stable learning.
Bayesian inference for spiking neural networks in speech processing. Explores weight uncertainty and loss landscape smoothing for temporal tasks.
Evidential Transformation Network: adapts pretrained models for post-hoc uncertainty estimation. Efficient alternative to ensembles/MC dropout for deployed models.
VOLTA: benchmark comparing uncertainty quantification methods for deep learning. Evaluates 10 UQ baselines across modalities and distribution shifts.
Game-theoretic analysis of creator incentives in multi-agent recommender systems. Cooperative game formulation for fair collaboration in bandit problems.
PRAGMA: foundation models for banking event sequences. Transformer-based architecture with self-supervised pretraining on financial transaction data.
Skip-Connected Policy Optimization (SKPO) for reinforcement learning with reasoning tasks. Improves upon GRPO by addressing high-variance advantage estimation.
EvoLen: evolution-guided tokenization approach for DNA language models. Addresses fundamental tokenization design challenges in biological sequence modeling.
Experience replay for LLM post-training RL formalizing optimal buffer design as trade-off between sample efficiency and data freshness.
Tensor decomposition method quantifying uncertainty in LLM-based multi-agent systems accounting for communication and role dependencies.
CLOVER framework for multi-agent RL cooperation conditioning value decomposition on realistic wireless communication graphs.
LottaLoRA training paradigm showing frozen random backbones with trained LoRA adapters recover 96-100% performance across diverse tasks.
Adaptive simulation experiment framework using pairwise comparisons to optimize LLM policies for operations management tasks.
Prompt optimization method decomposing reward variance into response and prompt variance to identify task amenability to optimization.
Multi-agent actor-critic reinforcement learning for disaster resilience controlling power, communication, and emergency response systems.
4-bit floating-point format (HiFloat4) for efficient language model pre-training on Ascend NPU hardware.
Guidance method for consistency models using joint flow distribution learning to enable classifier-free guidance without separate teacher model.
Training curriculum method for discrete flow-based image generation models to improve one-step sampling stability and quality.
Analysis of LoRA adapter spectral geometry to identify fine-tuning objectives and predict harmful model behavior in language models.
Safety steering mechanism for multimodal LLMs using dictionary-aligned concept control to prevent unsafe outputs without retraining.