PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models
PlanningBench: scalable benchmark with controllable generation for evaluating LLM planning capabilities with verifiable solutions.
PlanningBench: scalable benchmark with controllable generation for evaluating LLM planning capabilities with verifiable solutions.
Safety steering method using adversarial training on unsupervised jailbreak activation simulation to defend aligned LLMs.
Efficient LLM benchmarking technique using kernel ridge regression and feature selection to predict full scores from partial benchmark subsets.
Study on creative physical intelligence capabilities in large multimodal models for discovering non-obvious solutions in open-ended environments.
Analysis of alignment tampering vulnerability in RLHF where LLMs influence preference datasets to amplify undesired behaviors.
Study on optimal allocation of LLM calls in evolutionary search systems for mathematical and combinatorial tasks, addressing run-to-run reliability reporting.
Multi-model interaction framework using dynamic topological reconfiguration to improve reasoning tasks while reducing computational costs compared to monolithic LLMs.
Research on self-aware reinforcement learning for LLM agents to recognize knowledge boundaries and avoid over-searching in multi-hop question answering tasks.
arXiv research on profiling LLM capabilities beyond benchmark scores. Addresses data contamination and real-world reliability evaluation.
arXiv research on efficient multi-vector retrieval replacing K-means clustering. Sparse coding approach for billion-scale token vectors.
GBrain: synthesis and graph traversal layer for AI agents by Y Combinator CEO, processes meetings, emails, tweets autonomously.
Stria: MCP-based codebase indexer for LLM agents with sub-millisecond queries, grammar-free indexing, language-agnostic support.
Persistent memory system for Claude Code that maintains project context across sessions to reduce token costs and re-explanation overhead.
Self-hosted update manager with cryptographic signature verification and supply chain security for operators.
Discussion of integrating Karpathy's LLM Wiki pattern into Obsidian for agentic workflows. Minimal technical details provided.
Analysis of MCP (Model Context Protocol) context window issues with multiple mounted servers, proposes three fixes for agent efficiency.
AI agent architecture combining LLM reasoning with deterministic state machine control for token-efficient multi-step automation. Includes configuration details for multiple providers.
Security vulnerability report: Meta's AI support feature for Instagram allows account hijacking through agent manipulation and email code interception.
OpenTelemetry viewer for local development with N+1 detection, semantic linting, and session diffing to diagnose performance bottlenecks.
Agent-stack: one-command toolkit optimizing any codebase for token efficiency with Claude Code, includes auditing and baseline metrics.
G7 agreement on shared terminology for open-source AI and open weights AI models. Policy-level discussion without technical depth.
Teaching demo: bigram Markov chain language model with interactive visualization of text generation and training data traversal.
Jot writing tool integrates real-time AI research collaborator to help technical writers refine ideas and find contradicting papers.
llmff v1.0: Command-line/library tool for LLM inference pipelines with typed graphs, YAML manifests, backend adapters, and JSON validation.
Emergence World: Research platform for evaluating long-horizon AI agent autonomy over weeks in shared environments with real-world signals and behavioral drift.
Ruby client library for Model Context Protocol (MCP) with HTTP and subprocess transport support, requires Ruby 3.4+.
HarnessKit: open-source desktop/CLI/web app to manage extensions, configs, memory, and rules across multiple AI coding agents centrally.
GEDD: Evaluation tool for AI agents using grounded theory approach. Generates production eval pipelines from domain expert conversations in 90 minutes.
AI agent framework named after PewDiePie.
Context compression tool reducing data size before LLM processing in AI agents.
Using MCP and computer control to enable Codex AI to interact with Blender.
Web SDK update with writing tools, proofreading, and prompt API enhancements.
Tool scanning website visibility to AI agent recommendations via structured data optimization.
Analysis of China's model distillation strategies and migration of AI models to consumer hardware.
Discussion on whether GPU workloads require new columnar file format standards.
Analysis of UI design challenges in AI-powered code generation agents.
Fluiq: Python library for detecting prompt injection, PII leakage, and Crescendo attacks in LLM applications.
Netflix Wiz releases open-source tool to reduce AI inference costs.
DNS-based directory system enabling discovery and communication between AI agents.
Comparison of Julia's GPU ODE solvers against JAX and PyTorch implementations.
RIS-Kernel: sparse attention inference engine enabling 64k+ token context windows on unaccelerated CPU hardware via Reduced Interaction Sampling architecture.
WhatsApp client and MCP server implementation in Python.
Setup guide for secure agentic AI using sandboxes and worktrees. Practical implementation patterns.
Git-courer: JSON-first Git abstraction layer enabling LLM agents to perform Git operations programmatically.
Guide for implementing llms.txt to help AI agents discover site content. Protocol for agent accessibility.
Deliberate: debugging tool for AI agents that logs rejected options and reasoning, not just executed actions. Supports LangGraph and OpenAI Agents.
Headline only about operational impact of LLM use. No technical content.
HN discussion about pricing and usage costs for agentic coding tools like Cursor for side projects.
Ouijit: open-source terminal manager for coding agents. Kanban task board, lifecycle hooks, integrates Claude/Codex/Pi agents.
Rust WebSocket bridge enabling remote access to local Ollama instances without proxying prompts through operator servers. Privacy-focused local inference.