Scaling Pedagogical Pre-Training: From Optimal Mixing to 10B Tokens
Research on scaling pedagogical pre-training with optimal data mixing ratios, achieving improved perplexity on 10B token models with Sutra Framework.
Research on scaling pedagogical pre-training with optimal data mixing ratios, achieving improved perplexity on 10B token models with Sutra Framework.
Browser-based satellite imagery object detection using vision-language models with text prompts, outputs as GeoJSON.
Repository extending Shannon entropy with time-dependent learning observer for public research and audit, under restrictive source-available license.
Open-source CLI and GitHub Action for scanning MCP server security misconfigurations across configs, endpoints, and containers.
OpenGraviton open-source inference engine running 500B+ parameter LLMs locally via ternary quantization and dynamic sparsity.
OpenAI-compatible gateway with prompt compression, cost-aware routing, and budget enforcement for AI teams using Cursor.
Bevy-based tool for AI agents to build 3D explorable worlds and save them as reusable skills in skill directories.
Phi Guard is an open-source HIPAA PHI scanner for CI/CD pipelines that detects Protected Health Information in code repositories, complementing GitHub's Secret Scanning.
ANIP is a protocol standard enabling agents to reason before acting when calling APIs, with REST, GraphQL, and MCP adapters for agent-native interface design.
Hopsworks platform as runtime for coding agents like Claude Code, discusses agent infrastructure and data platform integration.
Case study of Claude Code building and testing a social application, demonstrating end-to-end test generation capabilities.
Community-maintained security baseline checklist for Model Context Protocol server deployments and AI agent infrastructure.
Method injecting structured domain knowledge into image generation prompts to reduce historical hallucinations in AI image models.
CLI tool converting MCP servers to command-line interfaces, reducing LLM context token usage by 96-99% through on-demand tool discovery.
entropick tool for replacing software PRNG with physical randomness in LLM token sampling, supports vLLM and Transformers.
AI agent hook for Claude Code providing visibility into accumulated work-in-progress like uncommitted changes and open PRs.
ADK-Rust announces roadmap to build Rust-native platform for building and deploying production AI agents across cloud, edge, and enterprise environments.
SandVault is a CLI tool for macOS that sandboxes AI agents and shell commands in isolated user accounts, providing lightweight alternative to VMs.
Differentiable 3D scene generation from language using vision language models for physically feasible embodied agent interactions.
Framework for deploying real-time AI services with autonomous agents across device-edge-cloud infrastructure with resource constraints.
Evaluation suite measuring whether reasoning models can control their chain-of-thought outputs, testing CoT controllability.
Medical imaging agents that self-discover skills and adapt tool invocation strategies through experience in clinical workflows.
Framework for evolving agent benchmarks with dynamic environments and changing toolsets to test robustness of LLM-powered agents.
Benchmark and co-evolving agents for verifying factuality in deep research reports generated by search-augmented LLM agents.
Multi-agent system using LLMs to automate product concept evaluation by analyzing and scoring concepts across criteria.
Open-source PyPDDLEngine for LLM-based task planning using PDDL simulation, enabling agents to sequence actions via tool calls.
Conversational Demand Response mechanism enables bidirectional aggregator-prosumer coordination through agentic AI in energy systems.
EpisTwin neuro-symbolic architecture grounds personal AI reasoning in knowledge graphs and user-centric data for holistic sensemaking.
SAHOO framework monitors alignment drift in recursive self-improving systems via Goal Drift Index and constraint preservation safeguards.
Schema-gated agentic AI framework enables conversational flexibility while enforcing deterministic, auditable scientific workflow execution.
Hybrid two-stage framework combines logical options pretraining with deep reinforcement learning for aligned agent training.
LLMs with structured prompts generate auxiliary lemmas for constraint solving with inductive definitions, outperforming SMT/CHC solvers.
Empirical analysis of human-in-the-loop principles in AI application development, documenting organizational challenges and operational guidance.
Companion artistic apparatus integrates drawing robot with LLMs for human-machine collaborative co-creation in visual storytelling.
Design study of systematic literature review tools identifies friction points in iterative scholarly work with AI-assisted exploration.
Traversal-as-Policy distills LLM agent execution logs into Gated Behavior Trees for safe, verifiable, and robust agentic control.
Survey of molecular representations for AI in chemistry from NLP perspective, comparing machine-readable and scientist-understandable formats.
Omni-C single dense Transformer encoder compresses heterogeneous multimodal inputs without separate expert encoders or routing overhead.
NGDBench unified benchmark for neural graph database capabilities across finance, medicine, and AI agent tooling domains.
VDCook self-evolving video data platform generates multimodal datasets for MLLMs via natural language queries with real retrieval and synthesis modules.
Survey of human-data interaction and visual analytics challenges with foundation models, LLMs, and VLMs for large-scale heterogeneous data analysis.
EigenData platform automates function-calling agent training data lifecycle through multi-agent orchestration for tool-use LLM synthesis and auditing.
Case-based reasoning approach for LLM-to-SQL translation in healthcare domain, improving RAG for EHR database queries.
Instruction-conditioned method combining imitation learning and reinforcement learning for robotic manipulation refinement.
Benchmark for evaluating self-evolving language agents' capability to autonomously create and adapt tools from task requirements.
Multi-modal generative framework for CAD design using differentiable parametric surfaces and unannotated 3D mesh data.
Multi-agent system enabling high-level autonomous control across diverse robot platforms via unified agentic interface.
SCOUT system combines 3D scene graphs with semantic reasoning for efficient interactive object search without slow LLM inference.
Framework testing stability and manipulability of LLM moral judgments using perturbation analysis on ethical dilemmas.
Multi-agent LLM framework with RAG for detecting hardware vulnerabilities in HDL designs, addressing knowledge gaps in security verification.