The Elegant Laminar Flow of Moroccan Tea [video]
Dojo: declarative testing engine in Go acting as transparent proxy to assert, mock, or AI-evaluate application behavior via HTTP and database interception.
Dojo: declarative testing engine in Go acting as transparent proxy to assert, mock, or AI-evaluate application behavior via HTTP and database interception.
Claude Opus used to generate Chrome V8 exploit chains through iterative prompting. Demonstrates LLM capability for security vulnerability discovery with detailed exploitation workflow.
Proposes agent.lock file concept for reproducible AI coding agent behavior. Discusses determinism and version control for agentic systems.
DFlash: block diffusion model for speculative decoding achieving 6× speedup over EAGLE-3. Parallel token drafting improves LLM inference efficiency.
MCP server for improving AI agent knowledge of specialized hardware (Chimera GPNPU). Addresses hallucination problems in domain-specific contexts.
AIPOCH: curated library of 420+ medical research skills for AI agents. Open source tool for domain-specific agent capabilities.
Sal Khan discusses why AI revolution in education hasn't materialized yet, citing low student adoption of Khanmigo AI tutoring chatbot.
Centrality visualization tool for observing Claude Code agent operations on codebases, showing file graphs and token consumption.
Guide for version controlling Claude Code IDE setup using git for syncing across machines.
Research paper on Charts-of-Thought method for enhancing LLM visualization literacy and understanding of data representations.
Spatial Atlas framework instantiating compute-grounded reasoning paradigm for spatial-aware research agents with A2A architecture.
DocSeeker multimodal LLM system for long document understanding with evidence grounding, addressing signal-to-noise and weak supervision challenges.
Systematic investigation of on-policy distillation dynamics in LLM post-training, identifying conditions for success and failure mechanisms.
Security analysis of federated learning for LLMs, investigating attack surfaces and defenses against malicious clients in open environments.
GUIDE framework for LLM-driven spacecraft operations using in-context learning to improve agent decisions across episodes without weight updates.
CodeTracer framework for debugging and tracing AI agent state transitions, error propagation, and tool orchestration in code agents.
Adaptive Memory Crystallization architecture enabling continual learning in autonomous AI agents without catastrophic forgetting.
Analysis of sequence-level reward learning in reinforcement learning for reasoning models, addressing gradient cancellation and learning efficiency.
Langevin Gradient Descent algorithm with generalization guarantees for hyperparameter tuning in convex regression via learning to learn.
Graph-based hierarchical reinforcement learning for automated co-design of thermodynamic cycle parameters.
Pareto-optimal offline reinforcement learning via Tchebysheff scalarization for multi-objective LLM alignment and optimization.
KV Packet enables context-independent key-value caching for LLMs without recomputation, improving inference latency.
Studies whether dimensionality reduction via random projections preserves landscape features for exploratory landscape analysis.
Context-dependent anomaly detection framework for multimodal data recognizing that anomalies depend on contextual factors.
Bias-corrected adaptive conformal inference for multi-horizon time series forecasting with distribution shift adaptation.
Counterfactual invariant prediction framework prevents shortcut learning in TCR-pMHC binding neural prediction models.
Binomial gradient-based meta-learning approach to reduce computational overhead in gradient-based meta-learning.
Twin-pass chain-of-thought ensembling method to improve confidence estimation reliability in telecommunications LLMs.
MOONSHOT framework for multi-objective one-shot pruning of vision and large language models without retraining.
Combines active learning and input denoising to improve robustness of neural operators against adversarial perturbations.
Multi-task LLM framework with LoRA fine-tuning for automated cancer staging and biomarker extraction from pathology reports.
Uses LLMs to enrich knowledge graphs for medical concept representation in EHR mining and clinical prediction tasks.
TabDistill method leverages tabular foundation models to identify feature interactions for generalized additive models on tabular data.
Orthogonal Backfill compression for LLM multi-agent systems reducing KV cache relay costs while preserving communication context and information.
BioTrain framework enabling sub-50mW on-device fine-tuning for MCU-based wearable edge AI on biosignals addressing domain shift and privacy.
Diffusion sequence models and Transformer meta-models for in-context robot dynamics learning addressing distributional shifts and real-time constraints.
Fine-grained non-determinism evaluation in diffusion language models showing dataset-level metrics mask run-to-run variations and condition sensitivity.
WIN-U: Woodbury-informed Newton method for machine unlearning in LLMs enabling 'right to be forgotten' without requiring retain set data.
Forward-only KL-based sensitivity analysis for mixed-precision quantization of hybrid SSM-Transformer LLMs targeting edge device deployment.
Proposes Chain of Uncertain Rewards method for designing reward functions in RL using LLMs, addressing inefficiencies in manual reward design and capturing intermediate uncertainties.
Research on SFT-GRPO data overlap as a post-training hyperparameter for Lean 4 autoformalization using Qwen3-8B, ablating overlap percentages (0%, 30%, 100%).
Analysis of multi-timescale PPO revealing surrogate hacking when fusing multi-scale signals, proposing representation-focused solutions.
Study of self-supervised learning and predictive representation learning comparing alignment and reconstruction approaches.
Confidence-based test-time voting mechanism for latent recurrent neural networks enabling test-time scaling without explicit energy functions.
DynamicGate MLP architecture permitting concurrent learning and inference by separating routing from representation parameters.
Parameter-efficient quantum multi-task learning with shared backbone and task-specific heads for quantum neural networks.
RL approach for radiology report generation using evidence-aware rewards and self-correcting preference learning for clinical alignment.
Analyzes reward hacking vulnerabilities in RLHF and alignment approaches for LLMs, examining mechanisms and emergent misalignment issues.
Bayesian mitigation strategy for safer AI agents using expanded subjective reward range to prevent reward hacking via risk aversion.
Studies how learning rates regulate catastrophic overtraining in LLM fine-tuning through catastrophic forgetting lens.