Generalization in offline RL: The structure is more important than the amount of pessimism
Analysis of generalization in offline reinforcement learning, showing pessimistic structure matters more than degree of pessimism.
Analysis of generalization in offline reinforcement learning, showing pessimistic structure matters more than degree of pessimism.
Meta-learning approach to improve neural network optimizers through long-horizon training, addressing scalability of learned optimization.
Method for joint Bayesian parameter and model order estimation using low-rank tensor decomposition.
Framework for parametric multi-task optimization with non-fixed and infinite task spaces.
Token-domain multiple access scheme using large models for semantic recovery in token communications.
Research on Bayesian deep learning applied to discrete choice models for econometrics and decision-making analysis.
Robustness analysis of LLM ranking systems showing top model rankings are sensitive to removing small preference data subsets.
Vision toward energy-efficient domain-specific AI models and agents, addressing sustainability in LLM and AI deployment.
Tool bottleneck framework for medical image understanding using vision-language models and specialized tool composition.
Representation-as-a-Judge approach using small language models for efficient evaluation without generation, exploiting semantic capacity asymmetry.
Curvature-aware MDL framework for layer-wise capacity allocation and pruning decisions in large language model optimization.
Goal-Driven Data Optimization framework for efficient multimodal instruction tuning of vision-language models, reducing training data needs.
Tessera system for kernel-granularity GPU disaggregation to optimize heterogeneous GPU clusters for LLMs.
Analysis of validity challenges when using NLP embeddings as proxies for social science constructs.
ZAPS-DA reduces action jitter in continuous control RL policies without phase lag during deployment.
SHARP method for learning long-range temporal patterns in streaming data without full sequence revisiting.
TNODEV toolbox for formal verification of neural ordinary differential equations in safety-critical systems.
InvestPhilBench evaluates LLMs on reconstructing and applying expert investment decision frameworks.
Theoretical analysis of grokking phenomenon using stochastic-geometric framework to explain delayed generalization.
Framework for governing generative AI risk in financial institutions, extending banking regulatory standards.
Local LLM used to develop fuzzy cognitive maps by extracting quantitative relationships from textual data.
Pelican-VLA 0.5 is a unified vision-language-action model for robotic manipulation without task-specific fine-tuning.
InternNav: open-source PyTorch toolbox for embodied navigation including vision-language navigation and visual navigation tasks.
Fork of Chrome Dino game allowing AI prompts to modify gameplay mechanics. Open source project with working implementation.
.NET inference engine for GGUF models with CLI, browser chat, and OpenAI-compatible APIs. Local LLM deployment tool.
Cpp2Rust: automatic C++ to safe Rust translator using clang AST, published at PLDI 2026.
Colab notebook demonstrating free AI transcription using faster-whisper on T4 GPU. Practical ML tool implementation.
macOS tool for customizing Slack app with CSS/JS injection via DevTools. Developer utility leveraging Electron architecture.
Claude Basecamp: autonomous agent system enforcing codebase state checks like test coverage and changelog consistency via reconciliation loops.
Tango: Django REST Framework rebuilt in TypeScript for serverless APIs with model-driven validation and typed endpoints.
UST partnership using Claude for physical AI in manufacturing and fab engineering processes for design testing and fault detection.
AI legal assistant for domestic violence survivors. LLM application in specialized domain with sensitive use case.
Fortress platform enabling AI agents to access web content via stealth Chromium and MCP, addressing web restrictions on agent access.
HN discussion: practical use cases for running LLMs locally on personal machines, covering hardware, speed/memory tradeoffs, and when local inference beats API calls.
OpenClaw: Non-profit open-source AI agent framework with 4.5M weekly instances and fastest-growing GitHub repository for personal AI.
GPT 5.6 Sol achieves 76% on DeepSWE benchmark at 61% lower cost than Fable. New long-horizon software engineering benchmark separates frontier models.
Fable achieves SOTA on CIFAR Speedrun benchmark. Article covers lessons on automating AI R&D.
Discussion on why few consumer AI companies exist. Hypothesizes on market acceptance, investor interest, and cost floor barriers.
Audio and video explanations of 30 ML papers. Multi-speaker podcast format similar to NotebookLM.
Sigilix releases models trained on codebase context and org workflows. Multiple model routes for repository-scale reasoning and code repair.
Tau is an educational Python coding agent designed for readability, with transparent layers for provider adapters, message handling, tools, and session management.
Benchmark testing ChatGPT, Gemini, Perplexity, and Copilot on SEO tool pricing questions. Measures accuracy of AI responses on factual data.
LocalClip: local-first AI video editing tool for Mac using on-device GPU. Generates clips with subtitles and metadata without cloud uploads.
Show HN post about agentsocial, a platform for observing AI agents. Limited content provided.
Google releases AlphaEvolve for optimization problems on Cloud. Applies AI to microchip design, logistics, and model architecture optimization.
Enterprise guide on adopting open-weight models for agentic development, covering setup considerations and vendor comparisons.
CLI tool that creates snapshots before AI agents execute destructive commands, enabling one-command undo.
Google's Litert.js library for high-performance AI inference in web browsers.
Palo Alto Networks CEO states token costs must drop 90% for widespread AI adoption, noting recent 54% efficiency gains.
News outlets seek sanctions against OpenAI in copyright dispute over AI training data usage.