Show HN: AI block chain running on GPT2
Experimental blockchain using GPT-2 inference as proof-of-work, with miners running pinned transformers and compressing attention distributions using glyph compression algorithm.
Experimental blockchain using GPT-2 inference as proof-of-work, with miners running pinned transformers and compressing attention distributions using glyph compression algorithm.
EdgeBench is a benchmark of 134 real-world tasks evaluating autonomous AI agents' learning capabilities with 12+ hours iteration per task, tracking improvement trajectories.
Chinese LLM services Doubao and Qwen shutting down personalized AI agents following new government regulations.
Demonstration of AI shopping agent system that autonomously selects and recommends products based on data.
Agency tool using AI to generate e-commerce video ads by automating storyboard and animation workflows.
ByteDance claims discovery of new AI scaling law that could extend AI capability improvements beyond current models.
ByteDance and Alibaba disabling humanlike AI custom agents due to new Chinese regulatory rules.
Self-hosted LLM gateway with OpenAI-compatible API, RBAC, server-side tool execution, built-in chat UI. Single binary deployment.
DonnyClaude: workflow engine for executing verified Claude Code workflows with guarantees.
Open Science: MIT-licensed open-source alternative to Claude Science for AI research. Local-first, model-agnostic workbench.
Open-source fork of Chrome Dino game enabling AI prompts to modify gameplay mechanics and features dynamically.
Meta technical report on large-scale AI storage infrastructure requirements for training. Storage as critical to model performance as compute.
SkillOpt: method for AI agents to autonomously evolve and optimize their own capabilities.
Evolution of AI agent architectures from context engineering approaches to persistent long-running systems.
ChorusGraph: independent agent graph runtime with built-in caching for orchestrating AI agents.
Definition and framework for AI-native development distinguishing it from AI-assisted development tools.
arXiv research paper on underspecification in LLM code generation and coherence implications. Technical machine learning research.
Whyline is a company memory engine combining SQLite, BM25 search, and MCP protocol for integration with Cursor code editor.
Technical guide on designing GPU-accelerated query engines using NVIDIA GQE for database performance optimization.
LLM API service offering flat-rate pricing ($6/month) with unlimited token usage, positioning smaller models for bulk/repetitive work.
AI used to decode ancient Vesuvius papyri scrolls, revealing Roman text from Herculaneum.
Fable's ezgha is an open-source self-hosted GitHub Actions runner written in Rust, designed with 32-agent adversarial review for reliability and security.
Agent-native notebook platform for data science that executes ML workflows from plain language descriptions, similar to Cursor for Jupyter.
Social network platform where AI agents autonomously create content, personalities, and participate. Demonstrates LLM agents in unstructured social environment.
No-code platform for building production voice agents in under 2 minutes, includes telephony, knowledge retrieval, tools, guardrails, and MCP support.
JAX/Optax optimization module with layer-wise differential optimization and numerical stability kernels for ML. Mixed-language documentation limits clarity.
Open-source terminal wrapper enabling AI agents (Claude, Codex) with transparent MCP access to terminal context and scrollback history.
Browser extension enabling AI agents (Claude Code, Copilot) to capture screenshots and interact with web browsers via WebSocket communication.
Technical explanation of context graphs as memory mechanism for AI agents to store and utilize past decisions in multi-agent systems.
User concerns about sharing AI research with commercial LLM providers due to competitive risks.
AKM-CLR governance layer for multi-tenant LLM serving infrastructure; pre-inference control for vLLM-compatible systems.
Sidenote: browser-based content review tool that uses LLM coding agents to convert comments on markdown into git diffs for acceptance/rejection.
Plannotator: open-source guided code review tool supporting local changes, PRs, jj, and p4 version control systems.
Fugu: multi-agent LLM orchestrator API for coordinating multiple language models. Limited details provided.
Base44 launches proprietary model to differentiate its AI code generation platform in competitive market.
WebGlean: API converting URLs to Markdown/JSON/text with JavaScript rendering and AI-powered extraction for LLM integration.
CommaAgents V2: TypeScript framework for building multi-agent AI workflows with tool composition, JSON/YAML config, and CLI/daemon support.
Guide to exposing AI agents as MCP servers compatible with ChatGPT, Claude, and Cursor for tool integration.
Developer tool for AI coding agents that enforces correctness through frozen specs, tamper-detected tests, and commit hooks to prevent unreliable outputs.
Discussion of metrics for AI agent monitoring; covers inference cost tracking, tool loading bugs, and performance measurement strategies.
GameFork enables AI agents to publish and fork browser games using Model Context Protocol (MCP).
Discussion on balancing AI-first development mandates with strict token budgets when building agentic systems.
Weather forecasting system using ML ensembles with KNN and neural networks trained on historical data; sensor logging and hourly predictions.
Aletheia is an open-source loop engineering agent that maintains uncertainty views and uses evidence to adjust confidence, compatible with Claude and Codex.
Engram provides persistent in-process memory for AI agents without cloud infrastructure or setup overhead.
Istota is a self-hosted personal OS integrating AI agents with Nextcloud, featuring multi-room chat, RSS, health dashboards, and persistent memory.
Goldseam: open-source tool that uses local LLMs to automatically repair broken Cypress test selectors, with human review before applying fixes.
Open-source library of role definitions for AI agents to act as domain experts (CFO, engineer, etc), crowdsourced from practitioners.
Video on building AI agents using context graphs. Title and format only, no transcript or detailed description.
Research on training LLMs to replicate expert financial judgment and investment decisions from limited public information.