Show HN: Real-time virtual try-on using hand gestures and live video diffusion
Real-time virtual try-on application using hand gestures and live video diffusion for dynamic retail interfaces.
Real-time virtual try-on application using hand gestures and live video diffusion for dynamic retail interfaces.
Fast local spreadsheet viewer in Rust using GPUI, built primarily with AI coding tools. Supports XLSX and Parquet formats.
Agent Bazaar framework for simulating and evaluating multi-agent LLM systems in economic marketplaces for stability and integrity.
Tool converting PDFs/EPUBs into queryable Claude Code skills, enabling cross-document question answering with efficient token usage.
Personal account of working on LLMs, security, and open source AI at Meta (2022-2026), including Llama and CodeLlama development.
Red teaming agents for LLM testing using adversarial techniques and tools like PyRIT, Garak, and Promptfoo frameworks.
Google Search update adding AI Mode with Gemini 3.5 Flash for conversational search experience similar to ChatGPT.
Stub post title only, no content provided.
Open source AI sales development representative tool for automating LinkedIn sequences and cold email outreach.
Guide on web scraping and converting unstructured HTML to clean structured JSON data without custom parsers.
Benchmark comparing Claude Code agent performance building TypeScript backends across five frameworks with same tasks and evaluation rubric.
LLM-powered text adventure game where players conjure objects and the AI generates functionality and properties dynamically.
High-performance Android device control CLI built for AI agents. Provides low-latency command execution via binary protocol and state mirroring.
Reference architecture using SQLite graph to preserve reasoning and context in AI-generated code beyond individual sessions.
Project providing structured documentation skills for AI coding agents, maintaining docs in sync with code through hierarchical dependency matrices.
Founder built AI orchestration platform (Meerkats.ai) reaching $3k MRR in 4 weeks. Case study on GTM and growth strategy.
Dust raises $40M Series B funding to scale multiplayer AI platform for human-agent collaboration.
DeepSeek V4 Pro and V4 Flash models now available via HPC-AI.COM API with competitive pricing for LLM applications.
Merlin Labs extends autonomous pilot system from military to commercial cargo aircraft operations.
Companion app for managing and switching between multiple Claude Code AI agent sessions.
LLM-backed CLI tool that analyzes code changes to determine review thoroughness needed.
Java-based toolkit for building stateful, server-rendered web applications without JavaScript.
Interactive tool to visualize LLM token generation speeds across different hardware platforms and models.
Benchmark coupling vision-language models with LLM-driven autonomous agents for aerial road-damage detection from UAV imagery.
Novel loss function for LLMs to improve numerical prediction in math and code generation tasks, addressing limitations of standard maximum likelihood estimation.
Study on scaling 3D perception models for autonomous driving, analyzing impact of model scale on multi-sensor fusion and spatial understanding.
Bayesian hierarchical model for identifying heterogeneous deterioration patterns in infrastructure equipment using causal discovery.
Theoretical work showing contradiction graphs determine VC dimension for binary concept classes.
Research on whether vision-language models understand 3D spatial scenes or only catalog objects.
DiffCodeGen: test-time scaling method for code generation using coverage-guided approach with reduced token consumption.
Research on hyperparameter prediction for model-based image denoising using oracle supervision transfers.
Research on metrics for evaluating uncertainty-augmented systems in automated decision-making contexts.
AgentAtlas provides unified benchmark beyond single-metric leaderboards for evaluating LLM agents across multiple dimensions including task success, tool validity, safety, and robustness.
Pseudo-Formalization method for automatic proof verification translating informal proofs into formal languages to enable AI mathematical reasoning evaluation.
Theoretical analysis of transfer learning sample complexity using optimal transport, addressing efficiency gains for LLMs and generative AI models.
Bandit algorithm for smooth graph functions applicable to online content-based recommendation systems where item ratings correlate with neighbors.
Matrix completion method for heterogeneous data with multiple group memberships, preserving subgroup-specific variation in recommendation and neuroscience applications.
Framework for multi-agent collaboration addressing concurrent codebase edits through state management to prevent silent conflicts and integration failures.
GPU-accelerated Mahjong simulator in JAX for reinforcement learning research on multi-player imperfect-information games with high-dimensional state spaces.
Analyzes how self-training on LLM outputs restructures language rather than flattening it, showing surface markers amplify while deep syntax diminishes across five model variants.
Backdoor attacks on LLMs via inference optimization/compilation exploiting numerical side effects; proposes attack framework.
Sparse-LowRank attention for diffusion transformers using 3D RoPE; addresses efficiency bottleneck in long-sequence video generation.
Study of AI peer reviewers on Nature papers with 45 expert scientists; evaluates capabilities, limitations, and credibility.
DIVE embedding compression via self-limiting gradients; reduces dimensionality for vector search systems without severe overfitting.
WebGPU backend for llama.cpp enabling memory-efficient, multi-precision LLM inference directly in browser.
Selective multimodal fusion for information extraction; decides when vision evidence is needed and which images to use.
Conflict-aware guidance for composing multiple constraints in diffusion/flow models without fine-tuning.
LLM-simulated experiments suffer from user drift; interventions cause unintended latent attribute shifts making results observational.
Framework for measuring information flow locality in hierarchical reasoning using sparse autoencoders and activation patching.
GPU efficiency metric (OFU) for AI workloads on HPC systems using performance counters; no application instrumentation needed.