Grok 4.5 Benchmark Results
Benchmark results for Grok 4.5 proprietary LLM showing performance metrics, context window (500k tokens), and multimodal capabilities.
Benchmark results for Grok 4.5 proprietary LLM showing performance metrics, context window (500k tokens), and multimodal capabilities.
Browser extension hiding AI chat history during screen sharing to prevent accidental exposure of conversation data.
Tool converting circuit descriptions and component assets into KiCad projects, enabling AI agents to work on hardware design without parsing complex file formats.
Opinion piece criticizing overuse of AI-generated content in business communications, discussing quality degradation from excessive output.
GLM 5.2 model successfully runs on consumer hardware with capabilities comparable to Claude/GPT while avoiding out-of-memory errors.
Probelock is a lockfile mechanism designed for LLM tool calling, enabling safer function invocation in LLM-based agents.
Former GitHub CEO launches Entire, a Git hosting network optimized for AI agents and code generation workflows.
Greppy extends grep with code-navigation subcommands optimized for AI agents exploring codebases.
Cybersecurity AI (CAI) Dataset released for training and evaluating ML models on security tasks.
Evaluation framework for founders launching AI-built applications; practical considerations for AI product launches.
Analysis of how AI models improve code rewriting economics for popular tech stacks due to training data prevalence and codebase context.
Tensorlake runs Docker images for Harbor on microVM sandboxes, passing Terminal-Bench 2.1 tasks with cold starts in seconds.
Compendium is a shared workspace tool designed for teams and AI agents collaboration.
Guide to removing bloat from Claude Code's system prompt, reducing token payload by tens of thousands per request through context optimization.
Title indicates research on indirect prompt injection attacks against RAG pipelines.
Production-assessed benchmark for code agents evaluating entire trajectory including instruction following, tool use, error recovery, and communication.
Theoretical analysis of in-context search as approximate inference over reasoning traces, studying sampling complexity of reflection-driven LLM reasoning.
Hybrid approach combining LLMs with agent-based modeling to enable real-time adaptive decision-making in large-scale individual interaction simulations.
Uses quantum processor as calibrated belief-update service for sequential POMDP belief updates in autonomous systems under partial observability.
Studies open-weight DeepSeek V3.2 model on ARC-AGI-1 benchmark using strict compute budgets without test-time scaling or fine-tuning.
ReAct-style agent combining LLM reasoning with SageMath CAS for verifiable feedback on research-level computational and experimental mathematics problems.
Analyzes token economics of enterprise agentic AI, arguing orchestration design (harness layer) is key lever against token inflation per task.
Identifies instruction leakage problem in goal-conditioned world models for spatial reasoning and proposes goal-free dynamics fix for true perception grounding.
Large Behavioral Model learns customer decision-making from retail transactions using Person-Environment formulation for grounded behavior modeling.
Studies how AI agents including LLMs learn social norms to improve human-AI coordination in dynamic interactions through implicit shared expectations.
Proposes relative-scale evaluation paradigm where AI models generate adversarial challenges to measure intelligence beyond human-saturating benchmarks.
Analysis of multi-agent LLM safety systems decomposing pipeline effects into three mechanisms: reframing harmful intent, planner refusal, and executor delegation.
ImagingBench: benchmark evaluating vision-language models and agentic AI on 20 computational imaging tasks across optics, signal processing, and inverse problems.
Framework for detecting logical inconsistencies in chain-of-thought reasoning of LLMs without requiring controlled interventions on evaluation transcripts.
Agents optimize static atomic tool actions into reusable Standard Operating Procedures to reduce reasoning overhead and failure rates.
PA-SciML: verification-first workflow for LLM agents discovering surrogate models in scientific ML with physics-based auditing.
MIRA-Math benchmark for mathematical reasoning where solvers must request one missing atomic fact to solve problems.
Agentic Data Environments framework for autonomous agents operating across files, APIs, applications, and system state with failure bounds.
Deterministic gates detect silent policy violations in tool-using LLM agents where forbidden state transitions execute successfully.
Shows biased reward judges silently disable skill retirement in self-evolving LLM agents, causing library drift below baseline performance.
SpaCellAgent: LLM-based multi-agent framework for autonomous trajectory inference analysis in spatial and single-cell transcriptomics.
Pyligent framework for training correction-aware reasoning in LLMs: validated search over partial solution chains with failure recovery.
Compares LLM-generated vs expert-written reusable skills for AI data scientists across data cleaning, SQL, statistical testing, and formatting tasks.
RL post-training enables Transformers to compose primitive skills into higher-level reasoning strategies beyond base model capabilities.
Survey of 1,250 papers on recursive self-improvement in AI systems: self-refinement, self-reward, self-play, and autonomous research loops.
SkillCenter: open-source library with 216,938 structured skills for autonomous AI agents across 24 domain bundles with source grounding.
Institutional red-teaming methodology for testing deployment rules in multi-agent AI systems with IABench-CA benchmark spanning 228 contexts.
Investigates whether model-free RL agents can identify price manipulation more effectively than model-based approaches in asset markets.
Analyzes memory poisoning attacks on LLM agents with long-term memory that access emails, calendars, and code repositories.
TriRoute jointly optimizes attention resolution, expert selection, and KV-cache allocation in language models to decouple quality from inference cost.
Mixture-of-experts approach for long-term forecasting addressing dataset-level distribution shifts via regime modeling.
Horizon-scanning study identifying security and privacy challenges in agentic AI systems.
Framework for optimizing diffusion model sampling via dynamic preference learning on schedules and guidance.
Open foundation model for wearable motion sensing with systematic study of pretraining and scaling.
LLM-guided approach for time-series forecasting in industrial processes using semantic metadata.