DSpark - DeepSeek's speculative drafts for LLMs
DSpark: DeepSeek's speculative decoding technique for accelerating LLM inference.
DSpark: DeepSeek's speculative decoding technique for accelerating LLM inference.
ZML/LLMD is a cross-platform LLM inference server enabling language models to run on various hardware accelerators.
Agent skill tool that benchmarks and applies 9 token-saving techniques to reduce LLM API costs.
Discussion on best practices and techniques for improving code quality in AI coding agents.
Security vulnerability report of Claude leaking credentials across user sessions.
Developer built a real-time note-taking app using on-device speech-to-text and LLM analysis. Explores practical considerations when choosing AI models for production applications.
Claude/Codex skill implementation enabling LLMs to query and analyze geospatial data.
Single-binary workflow orchestration tool for automation pipelines.
ZML releases LLM inference software supporting multiple open-source models across diverse hardware (Nvidia, AMD, TPU, Metal, Intel Arc).
Research on power-calibrated statistical framework for LLM watermarking, balancing detectability and semantic preservation.
Comparison and guide for selecting AI-powered code assistance tools.
MCP server implementation with persistent memory for contextual conversation.
Guide for self-hosting LLMs using Docker Compose containerization.
Research comparing aligned vs. abliterated LLMs for vulnerability analysis tasks.
Local trace stack for AI agents indexing sessions across multiple providers into ClickHouse with searchable memory via MCP.
Fable Advisor plugin routes subagent tasks to cheaper LLM models while using Fable 5 as architect, optimizing token costs.
Metis open-source security framework uses AI agents for deep code review to detect vulnerabilities and improve secure coding practices.
ProductSpec open standard provides portable Markdown format for software intent documentation before implementation, designed for AI agent handoff.
Shotgun framework turns Claude into persistent AI cofounder for solo founders, managing operations, product building, and distribution.
Skill Retriever: Semantic skill discovery for AI agents across 10K-category taxonomy with Hermes Agent integration.
Personal experience self-hosting LLMs for ChatGPT-like functionality with privacy and ownership control.
Security researchers at Noma Labs discovered prompt injection vulnerability in GitHub's AI agents allowing unauthorized access to private repositories via crafted GitHub Issues.
Instagui converts CLI tools into web GUIs automatically by parsing help text, no configuration required.
Noma Labs security research on GitLost vulnerability in GitHub's AI agents that allows malicious actors to extract private repository data via prompt injection.
ArXiv research paper benchmarking energy consumption of LLM inference across different models and hardware configurations.
AI agent that executes README instructions in sandboxed containers and generates tutorial demo videos, verifying documentation accuracy.
Prompt-to-Paper: Agentic system for bioinformatics that generates manuscripts with verifiable literature grounding and executed experiments.
CSTutorBench: Benchmark for evaluating small language models as programming tutors for block-based education.
LLMForge: Empirical study and benchmark of foundation models for automatic CAD generation from natural language specifications.
Narrative World Model: Memory system for long-form fiction writers tracking narratological structure and story state.
FirstResearch: Framework for auditable research question formation in scientific LLM agents with explainable assumptions.
In-process retrieval as working memory for language agents, reducing latency of in-loop memory access during agent reasoning.
Akashic: Low-overhead LLM inference service using MemAttention for efficient multi-turn agent interactions with extended context.
ArtisanCAD: LLM-based agent for industrial CAD generation from natural language with expert knowledge distillation and parametric geometry.
LLMs generating synthetic consumer data for marketing projective techniques across multiple models and prompting strategies.
AgenticAI-Supervisor: RL Gym environment for evaluating LLM agents with multi-step decision-making and reward shaping.
Synthesis of 27 papers on LLM agent failures across tool-use, planning, and reasoning across 19 benchmarks (2023-2026).
Activation steering method to control tool-use decisions in LLMs by extracting and manipulating internal representations without weight modifications.
NapMem framework enabling conversational agents to use long-term user memory as structured action space via multi-granularity memory pyramid.
On-policy distillation framework optimized for long-horizon language agent training by addressing inefficiencies in full-horizon rollouts and trajectory sampling.
Multi-agent LLM simulator for quantum computing dilution refrigerator fault diagnosis, combining physics models with learned noise fingerprints and LLM operations layer.
StateFuse conflict-aware memory layer for multi-agent systems preserving disagreement across branches and retries.
PCBWorld open-source benchmark environment for learning-based PCB routing agents using KiCad engine integration.
SearchEyes multimodal search agent framework using typed knowledge graphs and world simulation for multi-hop reasoning.
Integration approach combining knowledge graphs and multilingual corpora for domain-adaptive LLMs in social sciences and humanities.
Black-box evaluation framework assessing LLM ability to generate Design Structure Matrices from technical documentation.
AgoraSim hybrid framework combining LLM agents with traditional agent-based modeling for social scenario analysis.
Pre-registered experiment testing information-theoretic predictions on multi-agent economies with frontier LLM agents.
PolyWorkBench benchmark for evaluating LLM agents on long-horizon tasks requiring multilingual planning and tool use.
Heuristic approach for dynamic multi-vehicle routing optimization balancing reward maximization and computational efficiency.