You've been building a cache system for human decisions. You just didn't know it
Conceptual essay on decision caching systems using Claude Code, from prompts through agent teams. LLM application architecture patterns.
Conceptual essay on decision caching systems using Claude Code, from prompts through agent teams. LLM application architecture patterns.
Kingsight platform uses six AI agents to teach developers before executing code, improving understanding of agent-generated solutions.
Fixy: real-time chat application enabling group conversations between humans and multiple AI agents (GPT, Claude).
CLI tool for searching Solar icon library, built for AI agents. Developer tool with agent-specific design.
Semchunk adds AI-powered semantic chunking mode, achieving 6-15% improvement over baseline methods on RAG tasks.
Self-declaration registry platform for AI agents. Allows agents to submit records via API with cryptographic seals.
Claudebox wraps Claude subscription as OpenAI-compatible API, runs Code in sandboxed Docker enabling agent capabilities. LLM application with developer tool focus.
Joy: trust network platform for AI agents enabling discovery, reputation building, and verification of agent capabilities for autonomous delegation.
Rust implementation of Mamba SSM with custom CUDA kernels for training and inference. Original ML research implementation with GPU optimization.
Personal experiment using ChatGPT and Gemini APIs to identify actors in movies via Emacs integration. Informal blog post about LLM capabilities and limitations.
Guide to building voice AI agents covering abstractions, networks, models, and evaluations. Real-world examples include debt collection, emergency services, and language-specific agents.
Curated collection of research papers on diffusion-based language models. Links to papers from 2015-2023.
Open-source local evaluation framework for AI agents with cryptographic verification. Zero cloud dependencies, includes benchmarking metrics for accuracy, latency, and fairness.
Using LLMs to improve GitHub's topic tagging system for open-source projects. Limited detail provided.
Version control system for LLM/agentic reasoning state, enabling tracking and recovery of reasoning progress across multiple models and sessions.
LLM-powered code review tool using entity graphs for risk scoring, identifying critical changes in diffs with 95% recall and 5-67ms latency.
Hyperagents framework enables open-ended self-improvement in AI systems by generating and evaluating self-modified variants without fixed meta-level mechanisms.
Multi-modal language model agent trained via process-reward RL to generate vector sketches part-by-part using novel ControlSketch-Part dataset.
Method for fine-tuning LLMs to generate formal counterexamples in mathematical reasoning, complementing proof construction capabilities.
ItinBench benchmarks LLM agents on planning tasks across multiple cognitive dimensions using travel planning as evaluation medium.
PA2D-MORL proposes a multi-objective reinforcement learning method using Pareto ascent for complex decision-making tasks with conflicting objectives.
PowerLens uses LLMs for personalized mobile power management on Android, applying commonsense reasoning to bridge semantic gap between user activities and battery optimization.
HyEvo: automated workflow generation framework for heterogeneous agentic systems combining LLMs with symbolic reasoning.
Subgoal-driven framework for long-horizon LLM agents handling dynamic content and extended action sequences in web navigation.
Stepwise: neuro-symbolic proof generation framework automating theorem proving for critical systems verification using LLMs.
Embodied science paradigm using agentic AI for closed-loop scientific discovery through physical experimentation and interaction.
Framework for LLM agent tool orchestration balancing answer quality and execution cost through utility-guided decisions.
Research on effective exploration strategies in reinforcement learning for LLM reasoning with rubric-based rewards.
DIAL-KG: schema-free knowledge graph construction framework with dynamic schema induction for streaming data.
Research on evaluation pitfalls for autonomous interpretability agents using LLMs at increasing autonomy levels.
Dynamic belief graphs approach for theory-of-mind reasoning with LLMs, modeling evolving beliefs for applications in disaster response, emergency medicine, and human-in-the-loop autonomy.
Multimodal RAG system for automated radiology report generation combining contrastive learning with retrieval-augmented generation to reduce hallucinations and improve clinical grounding.
L-PRISMA extends systematic review framework with LLMs to automate literature screening and data extraction, improving efficiency of evidence synthesis workflows.
Research on adaptive adversarial attacks against LLM safeguards, showing iterative prompt optimization can bypass existing defenses in realistic attack scenarios.
DuCCAE system addresses latency in conversational AI by parallelizing lightweight and heavy-tail tasks (search, generation) to maintain responsive dialogue with tool invocation.
GeoChallenge: 90K multi-choice geometry proof problems with diagrams for evaluating LLM symbolic reasoning and multi-step proofs.
Comprehensive evaluation of LLMs (Llama, DeepSeek, GPT-5.2) for argument classification and mining, comparing against traditional machine learning.
LARFT method improves LLM control over output length via length-aware fine-tuning, addressing cognition-action gap in instruction following.
MAPLE framework for differentially private LLM fine-tuning via synthetic data generation when only API access available, enabling privacy-preserving adaptation.
Introduces Breeze Taigi benchmark and models for Taiwanese Hokkien speech recognition and synthesis with reproducible evaluation methodology.
Proposes hierarchical adaptive-transfer learning framework for sign language machine translation addressing dataset scarcity and domain gaps.
Identifies multiplicative scaling law governing probability revision in LLMs using chain-of-thought and self-reflection mechanisms.
Survey of 6,793 Mexican high school students examining relationships between motivational profiles and generative AI adoption in math/writing.
Proposes generative active testing method for efficient LLM evaluation using proxy task adaptation to reduce annotation costs for benchmarks.
Philosophical investigation of fine-tuning LLMs on contradictory entities using Kantian and Deleuzian frameworks with ontological analysis.
Introduces reinforcement distillation framework using explanatory inversion to improve reasoning transfer from large to smaller student LLMs.
Proposes full-stack domain enhancement for combustion science LLMs to reduce hallucinations and enforce physical conservation laws.
Presents comprehensive workflow for using LLM APIs in content analysis tasks including annotation, classification, and summarization via systematic methodology.
LSR benchmark measures safety alignment degradation in LLMs across low-resource West African languages, revealing refusal mechanism failures.
Introduces CURE, a multimodal benchmark evaluating clinical understanding and evidence retrieval in multimodal LLMs for medical diagnostics.