Open source alternative to brain by perplexity
Rudi: LLM memory system using causal graphs instead of full transcripts. Reduces API costs and context limits.
Rudi: LLM memory system using causal graphs instead of full transcripts. Reduces API costs and context limits.
Autonomy is self-directed AI agent framework with autonomous execution, risk-aware action gating, and outcome evaluation via approval policies.
C2PA Verify is open-source Android app using Content Credentials standard to verify photo authenticity and detect AI-generated images.
Speakora is AI voice generator for creators to produce natural narration for YouTube, product demos, and brand content.
Transformers.js v4.0.0 released with new WebGPU runtime in C++. JavaScript/WebGPU LLM inference library.
MCP server wrapping macOS Vision framework for on-device OCR, PDF text extraction, barcode detection, and image classification without cloud APIs or token costs.
Compression library reducing LLM token usage by 60-95% for tool outputs, logs, and RAG chunks.
Analysis of how AI is unbundling content creation from control in CMS systems, changing their role rather than replacing them. Discusses AI agents as alternative.
Lightweight 0.2B image inpainting model achieving 10B-level performance through systematic diffusion backbone optimization.
AgentArk is open-source self-hosted AI agent runtime. Build agents from prompts/tools, deploy as apps/automations, monitor actions, manage context. Full lifecycle management.
Open-source GitHub Action combining eight security scanners with unified output format.
Analysis of increasing architectural complexity in modern LLMs compared to earlier Transformer-based designs.
Critical perspective on LLM-generated incident reports, discussing risks and limitations while acknowledging toil reduction benefits.
GoodQ4All: offline video processing system with transcription, diarization, semantic search using local Whisper, Qdrant, and SQLite.
LodeDB: embedded vector database for local RAG with GPU acceleration, achieving 24k-50k queries/sec with sub-millisecond persistence.
Agentcard: virtual card infrastructure enabling AI agents to interact with real-world services like DoorDash with payment processing.
P34 framework addressing optimism bias in ML models trained on historical business decisions, preventing over-optimistic production deployments.
DeepGate's automated model search produces edge AI models 45× faster and using 11× less RAM than reference implementations.
DeepGate compiler achieves 3× RAM reduction and 2× speedup vs Google TFLM on microcontrollers.
Interactive notebook teaching GPU programming fundamentals using NUMBA for Python developers to build GPU kernels without prior GPU experience.
Analysis arguing smaller LLMs with intelligent routing outperform frontier models for knowledge workers, optimizing cost and performance tradeoffs.
Evaluation of local open-source LLMs for language translation tasks, comparing against Google Translate with methodology discussion.
Visual desktop app for composing multi-agent Claude Code workflows with drag-and-drop canvas interface.
PostgreSQL MCP server with 135 tools for Claude Desktop and MCP-compatible AI systems. Lock-free connection pooling, sub-10ms latency.
Web interface framework for AI agents to interact with rich UIs beyond text. Enables agents to manipulate tables, images, and complex layouts.
GPT-2-scale language model built from scratch in pure C/CUDA without ML frameworks. Includes hand-written tokenizer, training pipeline, and CUDA engine.
Reverse engineering Qualcomm NPU compiler to enable faster edge deployment of ML models on NPU hardware.
Model Context Protocol server for Uptime.com monitoring, enabling AI assistants to manage infrastructure checks and alerts.
Natural language interface for querying tabular data using LLMs and SQL, available for R and Python.
Sanity MCP server handles 3M+ agent tool calls; article describes gathering feedback from AI agents using a Model Context Protocol integration.
Folderly survey shows 70% of B2B sales teams use AI for outbound email but face 79% delivery failure rates as recipients detect AI content.
Beast is a governance layer for AI coding agents that enforces output contracts, repairs patches, and optimizes tool calls before filesystem writes.
Estonia issuing digital identities to AI agents for legal/regulatory recognition.
11-LLM consensus engine to detect hallucinations in AI outputs. Developer tool for improving LLM reliability.
CLI tool converting webpages to clean Markdown for AI agents. Open source utility with piping support and crawling capability.
Case study on switching AI tools mid-project causing productivity loss. Brief anecdote without technical depth.
Argument for local, on-premise AI systems vs cloud AI for privacy and reliability. Conceptual analysis of deployment approaches.
Guide on LLM prompting and interaction techniques. Incomplete content.
Article about practical LLM engineering and hands-on development practices. Incomplete content.
High-performance MCP server provides fast code intelligence for AI agents, indexing large repositories in seconds with tree-sitter AST parsing across 158 languages.
Redteam is an adversarial agent-pair harness for AI-assisted coding that uses two independent models to review and challenge code outputs before merge.
User tests Google's Gemma4 12B local LLM on 8GB GPU, finding recent model improvements focus on consumer hardware efficiency over scale.
Benchmark evaluating whether vision-language models understand 3D spatial layout beyond object recognition via occlusion and geometry tasks.
Pseudo-Formalization method automatically converts AI-generated mathematical proofs into verifiable formal language representations.
Multi-agent reinforcement learning approach for safe, superhuman autonomous racing coordination in dynamic shared spaces.
Light Interaction provides training-free inference acceleration for interactive video world models via optimized attention and caching.
UltraEP system optimizes expert parallelism training and inference for mixture-of-experts models with dynamic load balancing.
ACUTE Protocol uses language model internal activations to improve confidence calibration, reducing overconfidence and enhancing trustworthiness.
Evaluates video large language models' ability to provide real-time mistake correction guidance during instructional task execution.
Physics-informed Kolmogorov-Arnold networks with domain-specific adaptation for modeling axisymmetric pulsar magnetosphere dynamics.