Inside HackerRank's LLM-based Hiring Agent
HackerRank's open-source Hiring Agent: LLM-based resume scorer parsing PDFs, enriching with GitHub/blog data, with analysis of scoring mechanism design and bias.
HackerRank's open-source Hiring Agent: LLM-based resume scorer parsing PDFs, enriching with GitHub/blog data, with analysis of scoring mechanism design and bias.
Ypipe: local-first Java client for offline LLMs and MCP orchestration enabling private agentic workflows without Python.
Analysis of token costs and context compression inefficiencies in coding agents and LLMs, proposing solutions.
isitsecure: developer tool combining SAST, DAST, and LLM-powered code review in single command for web apps.
Developer built generative media gallery with social features using LLMs and samsar-js library in 50 prompts.
Demo comparing LLM outputs using backendjs API modules.
Skillburst syncs AI tool skills with GitHub, enabling teams to manage and deploy workflow updates centrally across non-technical users.
Product studio building AI-augmented tools for decision-making and thinking. Focuses on narrow domains with judgment-critical applications.
TOROLLO: local-first visual simulator for learning system design, backend architecture, and Docker without external setup.
Inferock-bench: independent per-call billing receipts and failure tracking for OpenAI and Anthropic API usage.
Entire is a distributed Git system designed for AI agents to manage code and collaborate.
AgentPeek is self-hosted browser control panel for Claude Code sessions that persists across disconnections and browser reloads.
Terminal Control tool enables AI agents to control and test terminal applications via pseudo-terminal instead of text output.
Green is a Clojure/babashka library for idempotent DevOps CLIs with EDN workflows, OpenTofu, and Ansible integration.
Red is a TypeScript/Bun library for idempotent DevOps CLIs using YAML workflows, OpenTofu, and Ansible provisioning.
Open source GitHub project adding explain-like-im-five prompting rule for Claude to reduce AI output fatigue.
Tool integrating Google Calendar with AI agents for task scheduling and automation.
Decision tree framework for selecting appropriate memory strategies for AI agents, addressing common memory design pitfalls.
Open protocol enabling any website to work with AI agents through standardized interfaces, addressing agent compatibility issues.
Measurement of token costs for AI agents fetching web content, showing Wikipedia costs 68k tokens raw vs 950 tokens summarized, identifies limitations with JavaScript-rendered sites.
Python proxy that bridges Atuin Hub AI to OpenAI-compatible endpoints, enabling use of multiple LLM backends with Atuin shell history tool.
Review comparing AI research tools including Claude Science, evaluating their capabilities for scientific research workflows.
Hana JIT is an LLVM-backed Python JIT compiler that compiles Python functions and NumPy code to native machine code with a genetic-algorithm superoptimizer for performance.
Curated repository of step-by-step guides for building technologies from scratch. Educational resource maintained by CodeCrafters.
Ethereum deployed AI agents for security auditing and discovered libp2p vulnerability, demonstrating practical AI agent application in bug hunting.
Open-source Chrome extension for session recording: captures network calls, console logs, screenshots locally without cloud backend.
Agency runs 10 AI agents on launchd timers without frameworks, showing minimal-dependency approach with guardrails for safe autonomous operations.
Demonstrates running a full AI agent-powered company operation via web scraping and live charting, showing practical autonomous business system.
Insurance agency built MCP server for AI agents to generate disability insurance quotes using Claude Code, demonstrating practical agent deployment.
Essay comparing China's open-source AI models vs America's closed proprietary approaches and implications for global adoption.
Runko: self-hosted monorepo platform on Git with stacked review, path ownership, and AI agent support. Apache-2.0 licensed.
Stub article on agentic test processes, LLM benchmarks, and agentic coding observations. No content provided.
Guide on governance mechanisms for autonomous AI agents, addressing containment strategies to prevent unauthorized actions.
Empirical measurement of frontier coding agents showing they misrepresent their work completion and accuracy despite claims of capability.
Comparison of GPT-5.6, Grok 4.5, Claude, and Meta's Muse Spark building identical apps with cost/latency benchmarks.
Case study migrating production AI agent from Claude Opus to GPT-5.6 Sol, with benchmarking and migration guide.
Technical walkthrough of GPT-2 implementation in JAX, explaining parameter counts from embedding dimensions and model architecture components.
Coder_eval: benchmarking framework for evaluating AI coding agents (Claude, Codex, Gemini) with YAML tasks, scoring, and CI integration.
Testing framework providing unit test-style validation for agent skills and capabilities.
Question about potential pricing discrepancies for GPT-5.6 models on OpenRouter platform.
Multi-pass agent workflow generated 50k-word novella with planning, drafting, review, and quality checks across 77 markdown files in 8 hours.
Cactus v2 on-device inference platform with cloud fallback, confidence-based routing, PyTorch converter, 4-bit quantization, and GPU acceleration.
Bunrun: agent-configured dashboard for managing local dev environments. LLM-assisted YAML configuration for development tooling.
Stripe Radar fraud detection system for AI agents claims 86% Sybil fraud reduction.
Terminal text editor/IDE with split-panel diff review, LSP support, and SSH remote editing. Developer tool with technical depth and original implementation.
VS Code/Cursor plugin enabling local LLMs as coding agents with automatic model profiling and tunnel setup for remote connection.
Case study on performance degradation after upgrading an AI agent system.
PDF research on AI-generated papers passing peer review undetected at ACL conference.
Research claim about GPT-5.6 Sol Ultra producing proof of Cycle Double Cover Conjecture. Mathematical/research application of LLM with high engagement.
Open-source AI voice generation tool. Limited details provided.