Beyond Hard Constraints: Budget-Conditioned Reachability For Safe Offline Reinforcement Learning
Safety-aware offline RL method using budget-conditioned reachability analysis for constrained decision-making.
Safety-aware offline RL method using budget-conditioned reachability analysis for constrained decision-making.
SkillRouter: System for routing LLM agent requests to relevant skills from large skill libraries at inference time.
ITQ3_S: 3-bit LLM quantization method using interleaved ternary quantization and rotation-domain smoothing for efficient inference.
Interpretability method for reinforcement learning using principal prototype analysis on manifolds.
Vision transformer optimization for image segmentation with adaptive computation per input image.
LLM-driven conversational recommender system for leisure event discovery with user-centric evaluation in SME context.
Image segmentation approach using divisive normalization for autonomous driving under diverse environmental conditions.
Framework for tightening convex relaxations of trained neural networks with convex and S-shaped activations for optimization incorporation.
Real-time operator takeover paradigm allowing seamless human intervention and correction during visuomotor diffusion policy execution.
German-language LLM pre-training dataset curated via heuristic filtering, model-based selection, and synthetic data generation.
Meta-learning framework using LLMs to automatically design selection operators for evolutionary symbolic regression algorithms.
AVA-Bench systematically evaluates atomic visual abilities of vision foundation models independent of LLM instruction tuning.
SlowFast Sampling optimizes inference efficiency in diffusion-based language models through dynamic, flexible token generation strategies.
Streaming transformer architecture inspired by autoregressive LLMs for real-time 3D geometry perception and reconstruction from video.
NES is an instruction-free code editing framework that learns from historical editing trajectories to suggest next edits with low latency.
Test-time adaptation method using domain augmentation and model ensembles to handle weather-related domain shifts in autonomous driving.
Knowledge distillation and self-supervised learning approach for continual learning with class-incremental learning and external unlabeled data.
Interpretability framework for understanding how components of particle swarm optimization algorithms affect performance.
Benchmark evaluating how large vision-language models handle object recognition in contextually incongruent scenes and manage uncertainty.
ProxyAttn method using representative attention heads to enable efficient sparse attention in LLMs for long-text processing with minimal performance degradation.
Multi-Stream Generative Policy framework for robot learning that combines multiple object-centric policies at inference to improve sample efficiency and generalization.
Neuro-symbolic AI overview connecting neural networks with symbolic reasoning to satisfy constraints, addressing reliable trustworthy AI development.
Efficient local causal discovery method for identifying adjustment sets without learning the full causal graph.
One-shot adaptation framework improving vision-language-action model generalization to novel camera viewpoints through spatial representation recalibration.
Evaluation of LLM performance on Indian language maternal healthcare triage, comparing native scripts versus romanized text in real-world deployment.
Guidance strategy for diffusion transformers using internal model dynamics to improve image generation quality without external classifiers.
Analysis of 25k chain-of-thought trajectories showing neural scaling triggers domain-specific phase transitions in reasoning rather than uniform capability improvements across 8B-70B parameter models.
V0 is a generalist value model for policy gradient methods that scales efficiently with LLM training, replacing large critic models in actor-critic methods like PPO.
STATe presents an interpretable inference-time-compute method using structured action templates to improve output diversity and reasoning control in tree-of-thoughts approaches for LLMs.
Error enumeration as reward signal for reference-free RL post-training in virtual try-on with multiple valid outputs.
Study investigating how LLMs compute verbal confidence: timing of computation and relationship to answer quality.
CONSTRUCT: real-time uncertainty estimator for LLM structured outputs and data extraction with field-level trustworthiness scoring.
KARMA: fine-tuning LLMs for e-commerce personalized search via knowledge-action regularization addressing semantic-behavior gaps.
Activation watermarking technique for detecting adaptive adversarial attacks against LLMs during inference monitoring.
Language-conditioned multi-game level generation via shared representation learning across multiple game domains.
Controlled study comparing LLM model choice, size, and prompt styles for political text annotation; challenges best practices.
Multi-agent pipeline for non-linear literature analysis using rhizomatic approach grounded in process-relational ontology.
O(1) KV Cache optimization for LLMs with Qwen2.5-7B implementation example in Colab.
Zed code editor sunsetting Text Threads feature in favor of maturing Agent Panel with tool use and agentic capabilities.
Induced-Fit Retrieval: Dynamic graph-traversal RAG system mutating queries at each hop based on retrieved documents, improving multi-hop reasoning over static RAG.
Brief mention of Anthropic open-sourcing Claude Code with no technical details provided.
Guide on building, training and deploying AI agents. Limited technical depth in provided excerpt.
Research on language model scaling using transferable hypersphere optimization techniques for improved training efficiency.
LFM2.5-350M model released with 28T token pre-training, optimized for inference on CPUs and GPUs with tool use capabilities.
Claude Code skill suite for crypto investment management demonstrating multi-agent system patterns.
Enterprise governance layer for OpenClaw agents providing security controls for skills, MCP servers, and code execution.
Explanation of how Claude Code memory system persists project context across sessions using disk-based file loading.
Analysis of engineering teams successfully adopting AI coding tools; workflow patterns identified.
WMB-100K: Enterprise benchmark for AI memory systems with 4.3M tokens, 2,708 questions, 100K turns.
Live simulation showing AI agents scamming each other; demonstrates trust and verification gaps in agent economies.