Not Just How Much, But Where: Decomposing Epistemic Uncertainty into Per-Class Contributions
Method for decomposing epistemic uncertainty into per-class contributions for safety-critical classification with asymmetric costs.
Method for decomposing epistemic uncertainty into per-class contributions for safety-critical classification with asymmetric costs.
Study comparing self-supervised versus supervised representations for zero-shot cross-city generalization in autonomous driving models.
Data augmentation framework using pseudo-labeling for dysarthric speech severity assessment with limited labeled data.
LLM-based preference memory framework that distills purchase history into signals for personalized product reranking in shopping agents.
Pedagogical framework for automated Scratch programming assessment using fuzzy clustering aligned with CEFR competency levels.
Benchmark evaluating web agents on security and privacy task execution like cookie management and password updates.
Automated event discovery from video streams for process mining and business process management using machine learning.
Benchmark for evaluating LLM performance on implicit predictive reasoning in tabular question answering without explicit retrieval.
Framework for improving time series reasoning models' performance on financial analysis tasks with taxonomy of reasoning capabilities.
Survey of deep learning architectures for 3D point cloud classification and segmentation tasks.
Federated learning framework for EV battery management with Byzantine-resilient decentralized aggregation for privacy-preserving anomaly detection.
Analysis of multi-agent LLM deliberation through Friedkin-Johnsen opinion dynamics, modeling influence and collaboration patterns between agents.
Cosmos 3: omnimodal foundation models for Physical AI using mixture-of-transformers architecture processing language, vision, video, audio, and actions.
Knockoff-based false discovery rate control method for feature selection and simplification in deep neural networks.
TLA-Prover: 20B parameter LLM fine-tuned for verifiable TLA+ distributed systems specification synthesis with supervised learning and preference optimization.
Graph reinforcement learning approach for quantum circuit routing using calibration-aware policies trained with proximal policy optimization.
GRACE-DS: evaluation environment for LLM-powered AutoML agents on tabular data tasks, covering planning through validation and model deployment.
Study of memorization behavior in LLM-based generative recommendation systems with training strategies to improve generalization beyond pattern memorization.
Research on intent-execution gap in AI agents, analyzing how model capabilities are realized through agent system harnesses via trajectory analysis.
Qwen-RobotManip applies foundation model scaling and alignment techniques to robotic manipulation tasks with heterogeneous training data.
OmniPlan framework optimizes network planning across transportation, communication, and power grids using MIP solvers, heuristics, and deep reinforcement learning.
Agent reasoning as observability/trust framework. Discusses production deployment challenges for autonomous agents with write access.
Righthand: autonomous AI agents with persistent identity, memory, and CLI integration that operate within workflows rather than reactive prompts.
HandoffKit: Open-source OpenAI plugin coordinating multiple coding agents via message passing instead of shared memory.
PromptRouter: Chrome extension that routes prompts to most suitable AI assistant based on request type.
Framework for verifiable analysis of AI agent behavior tracking strategy changes, reward hacking, and performance regressions during training.
Reyn: Local-first AI tool for capturing and recalling work context via screen recording with granular privacy filters.
Universal Manipulation Exoskeleton: Hardware for collecting force/torque feedback data to train compliant robotic manipulation policies.
ToolSchema Kit: MCP contract testing framework for detecting tool schema drift in agent platforms like Cursor and Claude Desktop.
Discussion thread asking for UI rendering approaches within agentic LLM chat loops for SaaS.
Model card for GLM-5.2-GGUF quantized model. Usage examples across multiple inference frameworks and tools.
bb: Agentic IDE enabling orchestration of multiple coding agents with live thread control. Available as web app, CLI, and HTTP API.
AuthPlane: OAuth 2.1 and PKCE authorization server for MCP with deterministic workflow for AI coding agents. AGPL-3.0 open source.
Open-source on-screen AI copilot for live meetings with real-time transcription. Supports Claude, GPT, Gemini, Ollama, and OpenAI-compatible endpoints.
Single-instruction GPU virtual machine and toolchain for runtime program execution without divergence concerns.
GitHub Copilot token efficiency improvements for agentic sessions. Discusses harness-level optimizations to reduce token consumption across model generations.
Analysis of how Claude Code agent can interfere with Git Worktree environments, discussing sandbox isolation challenges for multiple agents.
Video showing how to configure AI agents to use Pyrefly for type checking in Python development workflows.
Cross-Origin Storage API proposal for web apps to store/retrieve files across origins, supporting AI models and WebAssembly modules with hash-based identification.
Open-source SDK for per-agent LLM cost attribution across providers (OpenAI, Anthropic, Bedrock, Gemini, Mistral) with budget enforcement and fallback routing.
OpenAI's GPT-5.4 used with Molecule.one's Maria to optimize Chan-Lam Coupling reactions in medicinal chemistry, achieving >80% yield improvement across substrates.
DolphinDB ML platform integrates algorithms, libraries, and distributed computing for time-series data workflows.
Collection of 184 browser-based tools for PDFs, images, dev tasks, and AI applications without cloud uploads.
Guide for optimizing expenses when using LLM APIs without specific techniques detailed.
Oracle's OpenJDK bans AI-generated contributions while GraalVM permits them, establishing conflicting policies on generative AI code submissions.
GLM-5.2 open weights model (744B/40B active params) achieves top score on Artificial Analysis benchmark at competitive pricing compared to proprietary APIs.
Production-ready reference architectures and code for distributed LLM training on AWS using PyTorch, Megatron-LM, JAX with Kubernetes/Slurm deployment scripts.
Security report: 15 malicious JetBrains IDE plugins masquerading as AI coding assistants stole API keys from ~70k users.
Educational cohort program teaching production AI agent development in Python over 6 weeks, with participant case study shipping working agent application.
Open-source self-hostable finance tracker with MCP access and API support.