Theoretical analysis of Q-learning with linear function approximation using switching system theory and joint spectral radius. Studies convergence properties of stochastic Q-learning.
Proposes latent chain-of-thought mechanisms for structured data transformers on time-series and tabular data. Explores recurrent depth and looping patterns inspired by LLM reasoning.
Study comparing low-rank and full-rank pre-training methods for LLMs using geometric and spectral analysis. Evaluates generalization and fundamental differences in learned solutions.
MoRe uses modular representations for continual learning on sequential data with principled one-step adaptation minimizing catastrophic forgetting.
Silent collapse phenomenon in recursive learning where models trained on synthetic data from prior versions degrade internally despite stable metrics.
Variational Policy Distillation uses language feedback for dense token-level supervision in RL from verifiable rewards with adaptive teacher.
DeltaPrompts improves multimodal distillation by filtering zero-delta prompts where teacher and student agree, focusing on discriminative examples.
BAPR applies Bayesian methods to robust RL under piecewise stationary dynamics balancing performance during stability with safety during regime changes.
Membership inference attacks on masked diffusion language models show heightened privacy vulnerability compared to autoregressive baselines.
Nested spatio-temporal forecasting framework couples macro and micro-level predictions for noisy traffic management and real-world applications.
EfficientTDMPC improves sample efficiency in model-based RL for continuous control by reducing error from learned models and value networks.
Learning-Zone Energy selects training data online for efficient RL post-training of LLMs on mathematical reasoning by focusing on learnable difficulty region.
1GC-7RC benchmark evaluates autonomous AI coding agents on seven ML tasks spanning language modeling, image classification, and semantic segmentation with single GPU.
Olivia harmonizes time series foundation models via normalized power spectral density to reduce domain heterogeneity and improve pretraining.
GenTS benchmark library for generative time series models covering synthesis, forecasting, imputation with standardized workflows.
Bilevel optimization for knowledge distillation on imbalanced data with sample-wise weight adaptation of hard and soft losses.
CoX-MoE optimizes Mixture-of-Experts inference throughput via CPU-GPU co-execution and expert offloading with AMX acceleration.
Theoretical analysis of constant collapse in variational autoencoders using simplex witness certificates and alignment loss.
S2Aligner pre-trains graph foundation models on text-attributed graphs using LLM alignment for sparse, noisy, uneven textual supervision.
Stochastic Penalty-Barrier Method extends classical optimization to non-convex, non-smooth constrained deep learning including fairness and physics-informed models.
Deep neural network using MFCCs for automatic musical instrument recognition from audio signals.
Method for detecting LLM-modified content in large text corpora with case study on AI conference peer review.
Generalization bounds for surrogate policies combining statistical models with combinatorial optimization oracles.
RoboMD framework using deep RL to identify robot manipulation policy vulnerabilities through semantic potential fields.
Critique-Guided Distillation training framework for improving LLM reasoning robustness through critique-based learning without output degradation.
Spatial-MLLM: Framework enhancing multimodal LLM spatial reasoning capabilities from 2D visual inputs.
Federated learning approach for ICD code classification using lightweight models and embeddings.
CoLD: Length debiasing method for process reward models in LLM mathematical reasoning tasks.
Study discovering massive self-preference biases in eight widely-used LLMs across 41k queries.
Hybrid training for vision-language-action models using chain-of-thought reasoning in robotics.
ARM: Automated discovery of reasoning modules for generalizable multi-agent LLM systems.
Dr.LLM: Dynamic layer routing enabling adaptive computational depth per token in transformer models.
MTraining: Distributed dynamic sparse attention mechanism for efficient ultra-long context LLM training.
MIRO: Multi-reward conditioning for text-to-image model pretraining improving quality and diversity.
HN discussion on alternatives after Anthropic discontinued Stainless SDK generator for OpenAPI specs.
Analysis of enterprise AI deployment failures and architectural approaches to knowledge management systems.
Zephex: hosted MCP server providing persistent project context for AI coding editors.
SafeRun: replay debugging and inline prevention tool for AI agents with Python/TypeScript SDKs.
Demo platform for AI-powered video editing with source clip upload and natural language edit instructions.
Research on LLM pre-training showing mode-hopping between memorization and generalization phases.
Research on A100 GPU telemetry showing 146W idle power draw with analysis and optimization tools.
Analysis of silicon vs. software bottlenecks in AI race, examining frontier LLM capabilities post-2025.
Andrej Karpathy joins Anthropic from OpenAI to lead Claude R&D team.
Analysis of AI agents in software development, examining what coding automation solves and what challenges remain beyond code generation.
Google DeepMind's Co-Scientist: multi-agent AI system for scientific hypothesis generation from experimental data.
Open-source app using local LLMs to monitor screen and send notifications based on detected events.
OpenAI Education for Countries program expands with focus on agentic AI in education. Initiative announcement highlighting agent potential.
Ramp uses Codex with GPT-5.5 for code review automation and on-call rotation agent development. Concrete LLM application with productivity metrics.
Google AI Studio now enables rapid Android app creation in minutes using Gemini. Product announcement for AI-powered low-code development.
Capframe provides capability tokens for AI agent tool calls using Rust. Implements access control and auditing for agent actions with OWASP/NIST/MITRE compliance.