The HydroGym Reinforcement Learning Platform for Fluid Dynamics
HydroGym is a reinforcement learning benchmark platform for fluid dynamics control and modeling.
HydroGym is a reinforcement learning benchmark platform for fluid dynamics control and modeling.
White-Box Sensitivity Auditing uses steering vectors for interpretable LLM auditing beyond black-box testing.
Finite Difference Flow Optimization applies reinforcement learning to post-train diffusion models for improved image synthesis.
Medical image spatial grounding framework extends vision-language models to 3D anatomical structure localization in medical imaging.
T-QPM extends vision-language models for out-of-distribution detection and domain generalization under temporal distribution shifts.
Analysis of prompt sensitivity in LLMs, showing shared lexical task representations explain behavioral variability between instruction and example-based prompting.
Study of embedding-based defense failures in LLM multi-agent systems, showing how malicious agents can manipulate group decisions and decisions.
Stable reversible ODE solvers for diffusion model inversion enabling improved text-guided image editing without accumulated inversion error.
ABC-DFL: automated Byzantine-resilient decentralized federated learning framework for EV battery intelligence with privacy preservation.
AssetOpsBench study showing knowledge graphs significantly improve LLM agent accuracy for industrial asset operations reasoning over structured data.
RhinoVLA: Vision-Language-Action model optimized for real-time robotic manipulation on edge hardware by reducing visual and context token overhead.
DeXposure-Claw: agentic system using LLMs for DeFi risk supervision, routing decisions through structured evidence for regulator-aligned decision-making.
Mechanistic study of safety alignment in LLMs using Contrastive Logit Steering to isolate and understand the 'refusal direction' as a linear feature.
Analysis of hidden inference costs in quantized reasoning LLMs, showing low-bit models generate longer chains of thought despite correct answers.
Research on behavioral foundation models for zero-shot transfer in RL, enabling agents to generate optimal policies for any reward function without task-specific learning.
Research showing that LLM-style scaling laws (loss vs. model/dataset size) apply to sensor data, with implications for AI economics and emergent capabilities.
Local video processing system that transcribes, clips, edits, and generates captions and thumbnails from raw footage without cloud dependencies.
Sibyl is a self-hosted, scalable agent memory system built on SurrealDB enabling parallel AI coding agents to share context and coordinate work.
Claude Code users report silent deletion of conversation transcripts older than 30 days due to a non-obvious default setting.
Rigorix compiles natural language development tasks into deterministic DAGs for repeatable, auditable AI-assisted software engineering with policy constraints.
Agentic OS: proactive AI assistant for automating tasks, scheduling, and file management. Limited details provided.
Tutorial on self-hosting LLMs on NVIDIA Jetson Orin Nano for local inference without third-party APIs.
Anthropic launches Claude Science: specialized LLM workbench for scientific research. Announcement stub.
Guardians of the Agents: formal verification methods for AI workflows. Limited content.
llmaker: open-source platform for running complete LLM stack locally with models, vector DB, embeddings, retrieval, and agents. Single-command deployment.
Ovid is a tool that verifies AI coding agent features work by recording proof videos; integrates with pi agent and supports any-language projects.
OpenMontage is an open-source AI agent system that automates video production tasks including research, scripting, asset generation, editing, and composition from plain language descriptions.
Alma: local-first MCP server for AI agents to maintain persistent user context (name, preferences, values) across vendors without vendor lock-in.
RunInfra: automated model optimization and deployment. Benchmarks models, kernels, GPUs and generates deployment kits in 5 minutes.
Claude Sonnet 5 benchmarking: achieves 53 on Intelligence Index with higher cost-per-task than Opus 4.8.
Self-hosted memory system for AI agents with privacy-tiered architecture for persistent context management.
Open source repository demonstrating 98 AI architectures using Claude Haiku achieving 93% of Fable 5 quality at 1/125th cost with benchmarks and architecture details.
Tool for monitoring and predicting LLM API costs to prevent budget overruns in AI products including agents, copilots, and automation platforms.
TakoVM provides secure sandboxed code execution for AI agents with isolated Docker containers, ephemeral workspaces, job queues, and gVisor support.
Project using Claude AI with constrained optimization for automated research tasks, exploring AI productivity claims with evidence.
Government website redesign project using AI faces delays. Reports on policy implementation challenges.
AI agent that analyzes Sentry errors and generates GitHub pull requests with fixes. Integrates via OAuth and GitHub App.
Mimir: local-first encrypted memory system for AI agents as single Rust binary. Developer tool for agent state management.
Meta releases open source code for non-invasive brain-scanning system that reads sentences. Novel neurotechnology with open source component.
Open source macOS voice-to-text dictation tool using local AI models. Practical developer tool for speech recognition.
Bb IDE orchestrates coding agents with CLI for agents, shell scripts, and automation bots integrated into sidebar workflow.
Multi-head classifier system for detecting agent failures (looping, reasoning errors) using small LLMs with custom vLLM kernel for production cost and latency efficiency.
Periskop: product discovery MCP/API tool designed for AI agents. Developer tool for agent integration.
Analysis of prompt injection attacks framed as role confusion vulnerability in LLMs. Security research on LLM robustness.
Kage framework adds verification and freshness checking to Google's Open Knowledge Format (OKF) for agent memory management, ensuring proper memory creation and validation.
Claude Sonnet 5 benchmark results announced. Minimal content with no details provided.
Case study using coding AI agents to implement RonSQL support in RonDB database. Practical application of AI agents.
Research on training models to replicate expert judgment in financial tasks using machine learning.
OrgForge simulation framework for enterprise AI addressing hallucinations through maintaining boundaries between processes and generated text.
Open-source tracing tool for distributed LLM agents, monitoring token usage and linking to GitHub PRs/issues for development workflow tracking.