Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
Empirical study of LLM robustness to perturbations in chain-of-thought reasoning across five corruption types.
Empirical study of LLM robustness to perturbations in chain-of-thought reasoning across five corruption types.
ConFu: Speculative decoding technique improving LLM inference speed by enhancing draft model quality for token verification.
Multi-expert learning-to-defer framework addressing architectural limitations in expert selection and gradient distribution during training.
COMPOSITE-STEM: 70 expert-written benchmark tasks in physics, biology, chemistry for evaluating AI agent reasoning on scientific discovery.
The Amazing Agent Race benchmark with 1,400 directed acyclic graph tool-use puzzles for evaluating LLM agent navigation complexity.
Triadic Suffix Tokenization scheme improving LLM numerical reasoning by preserving digit structure and magnitude markers.
Korean-language multimodal benchmark with 3,466 questions covering nine disciplines and cultural context.
Post-transformer adapter technique to correct suppressed factual probabilities in aligned LLMs using 0.02% parameters.
Swiss government funding initiative for foundation model research and open science artifacts. €10M GPU hours available for core ML and applications.
Essay on headless APIs for personal AI agents to interact with services (passports, hotels, banking, shopping). Practical agent applications.
Auxx.ai: customer support CRM combining Attio and n8n automation. Self-built for managing support messages and customer data.
Agentjail: minimal Linux sandbox for executing untrusted code in agents and build systems. Rust core in production, TypeScript SDK/UI open-source.
Nyx: open-source autonomous testing harness for AI agents. Detects logic bugs, reasoning failures, edge cases, jailbreaks via red-teaming.
Discussion on security practices for AI agents using MCP tools, including data sanitization, tool output monitoring, and isolation strategies.
Analysis of Apple Silicon optimization for local LLM inference and AI agents. Hardware demand surge driven by on-device AI model capability.
Open-source theoretical implementation of Claude Mythos model using Recurrent-Depth Transformer architecture with looped recursion.
Hacker News discussion comparing AI/ML team database access challenges to traditional BI tool implementations, exploring read replicas and data governance.
Autoloom: lightweight autonomous AI agent (~1500 lines) built on tinyloom with state management, cron scheduling, TUI, and webhook server.
lmcli v0.5.0: minimalist CLI tool for LLM interaction written in Go with agentic tool-calling support.
Drawmode is an MCP server that generates Excalidraw architecture diagrams with automatic layout.
Analysis arguing LLM operational costs are sustainable and declining, countering common claims about economic viability.
GEPA prompt optimization improved Claude Haiku's bug-solving rate by 20% on new bugs.
Techniques using fake tool calls to prompt-engineer chat-tuned LLMs into base model behavior
Cardynal: No-code platform for deploying AI support agents with visual workflows, omnichannel inbox, and knowledge base.
AbleCredit case study on building human-centered AI systems for non-technical loan officers in India using WhatsApp/Instagram interfaces.
Open-source macOS app converting natural language to optimized SQL queries using LLM schema understanding
Vinkius Desktop: Native application providing unified MCP server management across multiple AI clients on a single machine.
Debate-style app orchestrating multiple LLMs (GPT, Grok, Claude) with scoring and verification
Infrawise: Azure cloud optimization tool using RAG backend and React frontend for AI-generated cost recommendations and resource analysis.
Philosophical article examining ethical implications of AI and robotics development over 70 years. Theoretical rather than technical.
Tool for automated penetration testing of LLM endpoints with real-time chat monitoring. Minimal description provided.
Free online book on building autonomous business systems with AI agents, including architectural patterns and working code examples.
Model-free RL approach for humanoid robots to learn stair climbing with safety constraints
Remen: iOS notes app with on-device LLM for natural language search and photo sync via iCloud. Privacy-first alternative to Google Keep.
Real-world comparison of Claude Opus 4.7 vs 4.6 performance metrics from live coding sessions
Data analysis of 543 autonomous coding hours using Claude Code agent showing productivity metrics: 14,926 prompts, 2,314 sessions, 165 releases.
Rensei: CLI enabling AI agents to generate 3D models iteratively from code with screenshot feedback, then 3D print. Includes skill file for agents.
Article title about LLM bug bounty landscape in 2026. No content provided.
Study comparing CNN-based OCR with vision language models, showing hybrid approach combining both architectures outperforms either alone.
RootCX: Open-source production infrastructure for internal AI-coded apps and agents. Includes database, auth, secrets, scheduling, file storage.
Passmark open-source Playwright library for AI regression testing, enables automated test validation.
Stub title about sandboxed AI agent orchestration platform; no content provided.
Analysis of 500 Show HN pages scoring them against 15 common AI design patterns to quantify sterile aesthetic in AI-generated landing pages.
Clerk tool auto-summarizes Claude Code sessions to Markdown format for documentation.
MCP Spine middleware proxy for Model Context Protocol with security, routing, token control, and compliance between LLM clients and tool servers.
iappyxOS: AI-powered on-device Android APK generator without cloud/servers; users describe apps to AI for local building and installation.
Hacker News discussion about spending more time conversing with AI than humans, reflecting on agentic coding tools and interactions.
PEAC: tool for verifiable API and MCP calls using portable signed records, supports Node.js, Go, Python with offline verification.
Git timeline visualization of Claude system prompts showing changes between versions; created by converting Anthropic's published prompts into commits.
Libredesk: open-source helpdesk alternative to Zendesk/Intercom built with Go backend and Vue frontend, fully free with no enterprise paywall.