Visual-TCAV: concept-based attribution method for explaining CNN predictions with saliency maps showing how concepts contribute to classifications.
Whisper-GPT: generative LLM for speech and music handling continuous audio representations and discrete tokens simultaneously.
Deep Convolutional Interpreter for Time Series (DCIts): interpretable architecture for nonlinear multivariate time series with factorized components.
Level set estimation algorithm with stopping criterion for efficiently identifying regions where expensive-to-evaluate functions exceed a threshold.
CleanPatrick: large-scale benchmark for image data cleaning using 496k annotations from medical workers on dermatology images.
Neural combinatorial optimization using latent guided sampling for NP-hard problems in logistics, manufacturing, and drug discovery.
Framework for cost-aware routing in text-to-image diffusion models to adaptively balance generation quality and computational cost per prompt.
Parallel ASR decoder using masked diffusion for non-autoregressive multilingual speech recognition with reduced latency.
Diffusion model modification using manifold hypothesis to generate clean samples from noisy training data via improved inference.
Reinforcement learning approach to improve LLM truthfulness by encouraging models to recognize uncertainty and abstain on out-of-distribution questions.
Training set sampling method using gradient-guided furthest point sampling for efficient data selection in molecular ML problems.
Theoretical analysis of oracle complexity for bilevel optimization with nonconvex upper-level and strongly convex lower-level problems.
Benchmark for evaluating vision-language models on complex exploratory visual reasoning tasks requiring multi-step chain-of-questions reasoning.
AI application that analyzes YouTube videos in real-time to detect hate speech vs peace speech and provides user feedback on content tone.
mlr3mbo: modular toolbox for Bayesian optimization in R supporting multi-objective, batch, and async parallelization.
Evaluates sycophancy failure mode in LLMs used for agentic financial applications and decision-making.
Preference-aligned memory construction for on-device RAG systems supporting personal AI agents with privacy.
Study of accumulated message effect on LLM judgments across 84k API calls to 12 models from 5 providers.
Research on emergent alignment in LLMs through persona selection hypothesis, studying how finetuning affects model behavior and alignment.
Building persistent cognitive architecture for LLM agents using Elixir and OTP concurrent programming framework.
Anthropic and GitHub used AI agents to port Git from C to Rust, passing full C Git test suite. Library-first, memory-safe implementation.
Open source Rust poker bot with async agents, regret minimization trees, and modern algorithms. Demonstrates multi-agent reinforcement learning.
Title only about runtime guards for AI agents. Insufficient content provided.
Technical approach using Lyapunov stability theory to detect when LLM agents become unstable or spiral. Novel application of control theory.
Master's thesis on hardware architectures for ultrafast inference and online learning using Kolmogorov-Arnold Networks on FPGAs.
Apple partners with Google to integrate advanced LLMs into Siri voice agent. Discussion of LLM-powered agent deployment.
Anthropic announces Claude Fable 5 and Mythos 5 pricing tiers ($10/$50). Brief pricing update without technical depth.
arXiv paper studying low diversity in LLM-generated stories, analyzing narrative patterns and character repetition in model outputs.
Technical analysis of Token-In-Token-Out principle in agentic RL, addressing training-inference mismatch in rollout evaluation.
LLM coding agents suffer from benchmark overfitting. Introduces DODO, production-grounded benchmarking method using live telemetry.
Lightweight coding agent harness written in Go that manages Claude model selection and installation with SHA-256 verification.
Python library for building knowledge repositories from multi-source data with unified query interface, designed for long-horizon AI agents.
Linux Foundation announces Tokenomics Foundation for AI token cost management and standards. Infrastructure/operations focused.
Claude Fable 5 model now available in GitHub Copilot for autonomous coding tasks, designed for long-horizon work with improved efficiency.
Yann LeCun discusses world models as foundation for next-generation AI systems in video format.
Zero-dependency tool that analyzes cost and ROI of Claude Code projects by parsing session transcripts and displaying metrics via web dashboard.
Nodea is an open-source AI canvas application for managing complex projects.
Lore is an LLM proxy tool designed to manage context and memory for coding agents.
AI-native project management tool designed for agentic workflows. Generates contextual tickets, Jira-compatible export.
Guardian Runtime is PyPI-available firewall for AI coding agents intercepting prompts/responses locally to prevent token leaks and API key exposure.
Browser-based text anonymizer for detecting/redacting PII before sending to LLMs. Processes locally via encrypted HTTPS with optional cloud sync.
ARL is open-source drift detection and model adaptation controller that learns from delayed labels to mitigate distribution shift without full retraining.
RoboCo is self-hosted open-source multi-agent software development system with 20 AI agents coordinating as virtual software company.
User feedback about GitHub Copilot pricing model consuming credits unpredictably when using agent mode and code review features.
VQAScore expanded to text-to-video evaluation using 20+ VLMs. Open-source metric/reward model replacing CLIPScore, adopted by DeepMind, NVIDIA, ByteDance.
Discussion thread: Senior engineer questioning skill degradation from heavy AI tool usage, concerns about understanding less while doing more.
Google Skills repository contains open-source agent skills for Google Cloud and products, supporting installation and contributions.
Technical case study on retry semantics in Genkit AI framework. Documents issue where retry middleware fires during cooldown windows from Anthropic API.
YC startup using CCTV and computer vision to automatically measure freight dimensions in LTL trucking terminals, replacing manual dimensioning.
Skilly is macOS menubar AI tutor using OpenAI's realtime API to watch screen and teach creative software with voice and UI pointing.