Show HN: Driggsby – securely chat with your finances
MCP-based financial data aggregator using Plaid for AI clients. Demonstrates practical MCP application for secure data integration.
MCP-based financial data aggregator using Plaid for AI clients. Demonstrates practical MCP application for secure data integration.
Memory architecture for AI agents using external shared storage rather than internal models. Title only, minimal detail.
Open-source long-term memory system for AI agents. Title only, lacks implementation details.
Security analysis of AI agent vulnerabilities. Title only, unclear technical depth.
Conceptual project exploring AI agents representing ecosystem interests and legal rights, combining agent design with environmental protection frameworks.
CLI benchmark for evaluating LLM function calling across 30 test cases. Supports cloud and local models for agent workflow testing.
Open-source tool detecting LLM hallucinations via hidden state analysis. Achieves 0.90+ ROC-AUC on Gemma/Llama with <1ms latency.
Technique for running multiple parallel AI coding agents simultaneously using git worktrees to achieve 2-3X productivity improvement over sequential execution.
Flock v0.7.0: Open-source DuckDB extension enabling LLM operators and RAG pipelines natively in SQL. Adds Anthropic/multi-provider support.
Open-source memory layer for AI agents. Title only, lacks technical implementation details.
Shell-based iterative coding approach for AI. Title only, insufficient detail provided.
Open-source autonomous agent runtime connecting AI to business systems (ERP, databases) via WhatsApp, Slack, Telegram with action capabilities.
Discussion on testing tools for MCP servers after Promptfoo acquisition. MCPSpec project for CI testing of Model Context Protocol.
MUP (Model UI Protocol) enables interactive UI components in LLM chat, allowing both users and agents to trigger functions. Includes PoC host and 9 example implementations.
Open-source AI agent designed to perform physics research tasks autonomously.
Framework for reliable AI agent development addressing hallucination and task drift. Structured protocol for production agent deployments.
Analysis of MCP dynamic tool registration feature. Argues MCP enables advanced agent capabilities beyond static tool definitions.
User question about AI tools for personal video editing. Discussion of limitations in current LLM video capabilities.
Performance comparison of Claude vs Calmkeep on 25-turn code and legal tasks. Shows 60%-85% code accuracy and 50%-100% legal accuracy.
Analysis of LLM competence zones for software engineering tasks. Framework for understanding model capabilities and limitations.
Benchmark study showing LLM code generation relies on memorization. Models score 90% on Python but 3.8% on esoteric languages.
Open-source voice-to-text tool with real-time speech cleaning and injection into any app. Customizable alternative to Whisper Flow.
ClickSay is a Chrome extension that captures UI context (selectors, styles, HTML, screenshots) and voice input for AI coding tools like Claude Code.
Security research showing AI agents can perform SIEM/EDR evasion, indicating organizations must assume adversaries will gain these LLM-powered capabilities.
Experience report using Lima for sandboxing AI coding agents (Claude Code, Codex) to enable autonomous operation with controlled permissions.
OpenAI releases GPT-5.4 mini and nano models optimized for coding and subagents with 2x faster inference and improved reasoning.
Rtk is a Rust CLI proxy reducing LLM token consumption 60-90% by filtering and compressing command outputs before context, with <10ms overhead.
Discussion thread with technical questions about LLM mechanics: token stopping, prompt continuation, and next-token prediction behavior.
Prototype using LLMs for autonomous assumed-breach penetration testing against Active Directory networks, demonstrating LLM capabilities in enterprise security contexts.
MarCognity-AI is an open-source framework analyzing LLM claim verification, finding 8-15% unverifiable claims. Decomposes responses and verifies against sources.
Primer on out-of-context reasoning in LLMs: when models reach conclusions requiring reasoning not present in context window, affecting generalization and alignment.
ModelSweep is a GUI-based benchmarking workbench for evaluating local LLMs on Ollama, enabling test suite building and comparative dashboards.
Llmgate is a lightweight Python wrapper supporting 21 LLM providers via YAML config with only 2 dependencies (httpx, pyyaml).
DataFlow is a low-code visual pipeline tool for generating, cleaning, and preparing high-quality LLM training datasets with flexible orchestration.
Web-based MCP server inspector tool for testing Claude MCP implementations with shareable URL configurations.
Online RL framework enabling LLM agents to evolve through retrospective dual intrinsic feedback and self-reflection. Code released.
Study investigating how humans attribute causal responsibility in AI-related harmful incidents involving agency, misuse, and misalignment.
Dual-path generative framework for real-time fraud detection in banking that balances low-latency anomaly detection with GDPR explainability requirements.
Benchmark evaluating zero-shot LLM prompting strategies for detecting security vulnerabilities in Solidity smart contracts.
Plan conditioning method improving diffusion language models' multi-step reasoning by prepending autoregressive plans to guide iterative denoising.
AI system automating document intelligence for UK planning authorities to navigate Data Protection Act and Planning Act compliance in large document volumes.
Multi-axis interpretable trust framework for detecting account hijacking using multidimensional criteria instead of single anomaly scores.
Safety framework for agentic AI systems that enforces deterministic pre-execution checks on real-world actions before API calls and transactions execute.
ManiBench benchmark for evaluating LLM code generation quality for Manim animations, testing syntactic hallucinations and visual-logic correctness.
Hierarchical fuzzy rule distillation framework for interpreting deep reinforcement learning policies in safety-critical applications.
Method integrating knowledge graph retrieval with LLMs to improve multi-hop reasoning and reduce hallucination through symbolic knowledge enhancement.
Agent-based framework for personalized harassment filtering in social networks using adaptive filtering agents that learn from user feedback.
Multi-agent LLM platform with deliberation-first orchestration for autonomous research automation combining tool use, reasoning, and code generation.
First-principles theory explaining the delay between memorization and generalization in neural network grokking with weight decay analysis.
LLM-based automated heuristic design system that dynamically evolves algorithms for non-stationary combinatorial optimization tasks.