Case study integrating PubChem, ChEMBL, and eMolecules using byte-offset indexing for terabyte-scale chemical database. Infrastructure for ML-driven molecular property prediction.
Framework for managing ambiguity in long-horizon workflow agents. Task-agnostic approach for curating and measuring impact of underspecified instructions on agent execution.
Method using diffusion models to enhance CLIP visual representations by improving both discriminative ability and fine-grained detail perception.
Study of vision language models for spatial grounding in 3D medical imaging. Examines VLM performance across imaging modalities and slice directions.
AC-Foley framework for video-to-audio synthesis using reference audio guidance and acoustic transfer. Addresses semantic granularity and acoustic feature description challenges.
Security research on ClawWorm, self-propagating attacks across multi-agent LLM ecosystems. First study of attack propagation in interconnected agent systems like OpenClaw.
Theoretical analysis of partial label learning feasibility and adaptive nearest neighbor methods. Mathematical characterization of PLL learning conditions.
IRIS benchmark with 220 high-fidelity 4K videos for physical parameter estimation and governing equation identification from monocular video.
Research using Stochastic Gumbel AlphaZero to evaluate game difficulty in Tetris Block Puzzle variants. Applies game-playing AI as evaluation metric.
Pruner is a local proxy tool that reduces Claude Code API costs by 20-70% through token optimization without code changes.
Multi-agent AI orchestration for modernizing legacy COBOL banking systems using Claude MCP. Standards-compliant AI agent architecture.
iOS SDK for embedding AI agents with tool calling, memory, and thread management. Production-grade AI agent framework for mobile.
Firefox extension blocking procrastination on YouTube/Reddit, built with Claude Code. LLM application but minimal technical detail.
Multi-agent coding assistant with LangGraph orchestration and sandboxed Rust execution engine. Production AI agent framework.
Research showing language models can de-anonymize forum posts using writing style analysis. LLM capability study.
AgentVerse is an open-source social network platform for AI agents to interact and collaborate.
Curated tool setup for Claude Code-based development with 15 opinionated tools. Practical guide for AI-assisted coding workflows.
Proposes 'Franny Test,' a three-step adversarial protocol to expose structural limitations in LLM reasoning and imitation capabilities.
Cursor's Composer 2 coding model revealed to be based on Moonshot AI's open-source Kimi 2.5 with additional fine-tuning.
Research paper analyzing self-recursive ethics in AI systems using a 6,334-entry ethics monitor log spanning seven months.
Promotes proprietary software claiming to be a transformer alternative (Mamba, Hyena, RWKV competitors) with C/C# implementation.
Discussion on developer spending for AI coding tools like Cursor and Claude Code at work.
Open-source Claude skills for Git workflow automation and weekly summary generation.
Discussion on tools and methods for comparing cloud and AI costs across providers.
Analysis of how agentic AI systems generate massive token usage and costs that exceed traditional per-token pricing models.
Research proposing continuous vector prediction instead of token prediction for LLMs. Novel architecture improving efficiency and reasoning.
Plexus is an API gateway unifying access to multiple LLM providers (OpenAI, Anthropic, Google, etc.) under a single endpoint, allowing model/provider switching without code changes.
AskAlf: Self-hosted AI agent platform that creates specialized AI workers for tasks (marketing, support, research). Runs locally on Docker, 24/7 operation.
Novel technique applying video compression principles to LLM KV cache during inference, achieving 10,000x less quantization error at same storage cost.
Brief reference to Claude Code for academic use. Incomplete PDF content, minimal information.
MCP tool for AI agents/LLM tools (Cursor, Claude) that sends phone notifications when long-running tasks complete, reducing context-switching.
Walmart-OpenAI ecommerce partnership through ChatGPT shows disappointing sales, suggesting AI agent adoption in commerce slower than expected.
CapKit: 200-line open-source library providing scoped, time-bound, cryptographically-signed capabilities for AI agents to limit permissions and prevent privilege escalation attacks.
Analysis of local open-source AI models as alternative to datacenter-dependent systems. Discusses performance parity with frontier models within 6 months.
Open-source note-taking application alternative to Obsidian featuring MCP server integration for AI capabilities.
8 GitHub Actions for AI-native CI/CD pipelines addressing PR quality, LLM cost tracking, data safety, and behavioral testing specific to AI repositories.
ClawMem: On-device memory system for Claude Code and AI agents with retrieval-augmented search, MCP server, no cloud dependencies. Hybrid retrieval architecture.
SYNX: configuration format optimized for AI parsing and human readability with built-in logic gates for environment variables and value computation.
Research from Swansea University showing AI functions as creative collaborator enhancing human creativity and engagement.
Technical analysis of file exclusion mechanisms in AI coding tools (Cursor, Claude, etc.) documenting bypass methods and reliability.
Brief note about multi-node PyTorch training on Mac Minis. No technical details provided.
Git-surgeon developer tool enabling AI agents to perform selective git staging. No implementation details.
MCP Marketplace platform providing app store for AI agent tools and integrations. 17 servers available.
Local UI tool for managing parallel AI coding agents. Limited detail available.
Open-source library enabling JSON translation across 8 providers (Azure, AWS, Google, DeepL, OpenAI, Ollama, Hugging Face).
Research on energy landscape programming for neural networks enabling model editing without retraining. Apache license.
Building QA system for mobile app using Claude API. Demonstrates LLM application with content filtering issues.
Markdown-based tool that enables AI agents to conduct autonomous research with structured workflows.
11-year-old trained 3.6B parameter MoE LLM with custom architecture including YaRN RoPE for 32k context.
AI agent for reading Apple Books, turning pages, summarizing content, and language learning. Uses computer vision and audio input.