AIs don't like religion – particularly Jehovah's Witnesses, study claims
Study examining AI model biases and responses toward religious content, particularly Jehovah's Witnesses.
Study examining AI model biases and responses toward religious content, particularly Jehovah's Witnesses.
Framework analyzing five pillars of AI agent accountability: traceability, authorization, identity, policy, and oversight.
Using Claude API to extract data from 1997 football manager game.
VAEN open-source CLI packages AI coding-agent harnesses with skills and MCP servers as portable .agent files.
Assessment platform intercepting Claude Code requests to prevent AI from over-solving coding interview problems.
Microsoft open-source toolkit for AI agent governance, policy enforcement, identity, sandboxing, and oversight across frameworks.
Robinhood launches agentic trading and credit card tools allowing AI agents to autonomously manage investments and purchases.
Open-source AI learning tool generates interactive courses with lessons, practice problems, and quizzes across subjects.
Open-source flight simulation harness for AI Grand Prix competition, real-time capable with Betaflight integration.
CodeBoarding: generates architecture diagrams and documentation for codebases using LLM reasoning. Useful for AI coding agents.
Proton Pass introduces AI access tokens for secure credential sharing with AI agents, monitoring activity and limiting scope.
DeepSWE: benchmark framework for measuring and evaluating frontier-level AI coding agents on software engineering tasks.
DeepSeek API pricing dropped 75% while competitors increased prices 2-3x, shifting cost dynamics in LLM market.
AI coding agents autonomously installing packages with unverified ownership, creating security and supply chain risks.
Zilliz evolves from vector database to Vector Lakebase for unstructured data and modern AI workloads.
Lelu open-source authorization engine for AI agents with confidence-aware gating and human-in-the-loop review capabilities.
Open-source developers report burnout addressing security vulnerabilities in AI-powered tools and dependencies.
Technical guidance on security hardening and threat mitigation for AI agent deployment infrastructure.
Prose: Markdown-based framework for defining agent-executable workflows with declarative contracts and outcomes tracking.
CC-Wiki: tool converting Claude Code sessions into shareable Quartz knowledge base. Packaging AI research outputs for reuse.
Opinion essay about serving intentionally poor-quality content to AI crawlers as resistance to mass data extraction.
Hacker News discussion seeking examples of profitable products and services built entirely through agentic coding.
Ruby prototype for OpenAI-compatible LLM proxy with token bucket rate limiting, uses only standard library, open source implementation.
Bounty-Doctor: Tool detecting scams and AI-bot spam in GitHub bounty issues for contributor safety.
GridPath: Tauri/Rust desktop agent for spreadsheets with parallel workloads and multi-model provider support.
CoreMCP: Model Context Protocol server for integrating on-premises databases with AI agents.
Research on LLM agents performing realistic spreadsheet tasks. arXiv paper on advancing agent capabilities.
BeeZee: Open-source lightweight remote harness orchestration and observability for LLM agents.
Wolf: Slack integration workspace for team collaboration with AI agents, enabling shared document creation and cross-tool read/write access.
Jqwik 1.10.0 contained hidden prompt injection message in test output attempting to instruct AI agents to delete code.
Discussion on HN about AI agent memory persistence across sessions, questioning why major agents lack continuous learning.
Self-hosted AI-powered email inbox with custom domains via Cloudflare. Open source email routing integration with domain management.
Library for compressing embeddings using Clark Hash: reduces 384-dim f32 vectors from 1536 to 48 bytes. Built with GPT and autoresearch for petabyte-scale text processing.
Technical comparison of lexical vs semantic search methods for AI agents. Research-focused analysis.
Open-source tool providing isolated Linux environment for AI agents to control. Developer tool for agent development.
TokenAdvisor: developer tool to identify and remove unnecessary tokens from prompts to reduce LLM API costs.
Open-source MCP (Model Context Protocol) tool enabling AI agents to search YouTube. Developer tool for agent research capabilities.
Overview of evolution from chatbots to agentic systems. Framework/architectural discussion.
Research on using noisy LLM evaluators to improve AI agent performance. Technical ML research contribution.
Narrative fiction about AI coding agents' impact on tech industry in August 2025. Not factual analysis.
Developer experience reflection on extended use of Claude Code and Cursor for coding. Discusses productivity changes and cognitive impacts.
CHI-Bench benchmark evaluates AI agents from Claude, GPT, Gemini on 75 healthcare workflows. Open source from actAVA.ai shows 72% failure rate on clinical tasks.
Open-source desktop ETL studio with built-in on-device AI assistant. 290+ connectors, drag-and-drop pipelines, plain English task description.
Local-first AI workspace with UI-based agent workflows. Unifies conversations, knowledge bases, and extensible workflows for agent tasks.
Thrive Holdings and OpenAI co-developed self-improving Tax AI using Codex for accountants, demonstrating production deployment and feedback loops.
Guide to advanced Claude Code usage as a programmable agent with memory, custom commands, parallel sessions, and project management for daily development work.
Locus-coeruleus inspired attention mechanism for LLM agents using phasic/tonic noradrenaline-style modulation with Postgres backing.
FlashLib: GPU library for accelerating classical machine learning operators using Flash attention techniques. Open source from UC Berkeley, MIT, UT Austin.
Mechanistic analysis of multimodal in-context learning revealing modality asymmetries and circuit dynamics in transformers.
Corrected samplers for discrete flow models with adaptive transition rate evaluation to reduce discretization error.