ARGUS: Adaptive Rotation-Invariant Geometric Unsupervised System
ARGUS detects distributional drift in high-dimensional data streams using local statistics over fixed spatial partitions of data manifold.
ARGUS detects distributional drift in high-dimensional data streams using local statistics over fixed spatial partitions of data manifold.
Stratified hazard sampling reduces variance in discrete diffusion/flow models by optimizing event scheduling in CTMC/DTMC processes.
PROMA: reference-free proximal policy method for LLM training that controls KL divergence via gradient projection without reference model.
OPO: theoretical framework for LLM alignment using constrained proximal policy optimization with work-dissipation principle and chi-square geometry.
Instant Retrospect Action algorithm improves policy exploitation in online RL through Q-network representation learning.
Green-NAS multi-objective neural architecture search optimizes weather forecasting models for efficiency and carbon footprint.
MinPV Principle minimizes path variance in score-based models to improve accuracy and stability.
Analyzes role of iterative computation in RL, showing policies benefit from additional compute beyond fixed parameters.
Algorithm achieving simultaneous optimal static and dynamic regret in adversarial multi-armed bandits.
Horizon Imagination improves efficiency of diffusion-based world models for RL by denoising multiple future observations.
Systematic evaluation of chemical language model scaling on molecular property prediction downstream tasks.
Recovery-based shielding framework integrates Gaussian process models with RL for provably safe control in continuous systems.
Online GPU energy optimization using bandit algorithms to reduce power consumption in HPC systems.
Extends linear bandits theory beyond inner product spaces using optimal transport for recommendation and clinical systems.
Evaluates multimodal LLMs and vision-language models for time series anomaly detection in systems monitoring.
MINT framework aligns LLMs with biomedical knowledge using preference optimization on multimodal data.
Cadrille uses reinforcement learning for multi-modal CAD reconstruction from point clouds, images, and text inputs.
AI agent framework integrating LLMs with Lean formal proof assistant for automated theorem proving.
Scaling long chain-of-thought reasoning in LLMs using NP-hard graph problems for cost-effective training.
Autonomous AI agents for physics data analysis using machine learning in particle physics research.
RLHF approaches for improving LLM-based UI generation using designer feedback and rationale.
Vision-language model enhancement framework for improved disaster assessment image descriptions.
Safety vulnerability analysis of diffusion language models and mitigation strategies for jailbreak attacks.
Method for scaling parallel LLM inference by enabling interdependent token generation across multiple responses.
arXiv paper extending KernelSHAP with interaction-informed polynomial regression for efficient Shapley value approximation.
arXiv paper on FlowSteer, an end-to-end reinforcement learning framework for interactive agentic workflow orchestration.
arXiv paper on Grappa, a distributed GNN training framework using gradient-only communication for scalability.
arXiv paper on online fine-tuning pretrained policies for autonomous driving using recurrent reinforcement learning.
arXiv paper on robust simulation-based inference using generalized Bayesian inference and neural network approximation.
arXiv paper on curriculum learning and pseudo-labeling for multi-label Arabic dialect identification.
arXiv paper on causal constraints in neural emulators of turbulent systems using response theory and score matching.
arXiv paper demonstrating tabular foundation models effectively learn association rules for knowledge discovery in tabular data.
Report that US Pentagon may require contractors to certify non-use of Anthropic's Claude API.
Analysis of how AI agent adoption disrupts seat-based SaaS pricing models and software business economics.
Security-audited directory of 458+ MCP skills and integrations for AI agents; graded on adoption and audit scores.
Business insight that 45% of support leaders are moving beyond generic AI chatbots to specialized implementations.
MCP server converting Reddit sentiment into options trading signals via Claude; 9-stage pipeline achieving 52% win rate on live trades.
Agent Audit Kit v0.1 provides deterministic replay and stress testing for LLM agents. Minimal details in title-only post.
DevDay CLI tool aggregates AI coding session data from multiple tools (Cursor, Claude Code, OpenCode), syncs with git, generates standup summaries locally.
Independent AI leaderboard with custom benchmarks testing real end-user and developer scenarios beyond saturated MMLU/HumanEval tests.
Guide explaining shift from chatbot interactions to agentic AI usage where models autonomously accomplish tasks with tools.
Conduit: Swift library providing unified interface across multiple AI providers (cloud and on-device). Actor-based architecture prevents vendor lock-in with one-line provider switching.
AgentForce: 15KB Python multi-LLM orchestrator replacing LangChain. 6.5x faster latency, 75% memory reduction, 1000x smaller footprint, 89% LLM cost reduction via Redis caching.
Incomplete submission title only, no content provided.
Edge-Veda: Managed on-device Flutter runtime for LLMs (text, vision, speech, RAG) with stability guarantees and privacy. Supports Llama 3.2, Qwen, SmolVLM models.
GreedyPhrase tokenizer achieves 1.21x better compression than GPT-4o tiktoken with 65K vocabulary and 6x faster throughput.
Guide covering AI agent design principles, patterns, and techniques including prompt engineering, specialist coordination, and iteration.
Command sandboxing solution designed for safe execution of code in AI agent environments.
First hackathon event for agent skills organized by SkillsBench authors.
First-person account of user operating OpenClaw agent over two weeks, describing agent capabilities and behaviors observed during extended use.