Show HN: Fuzzy-matching messy job board data against the UK Gov Visa Registry
Fuzzy-matching tool for matching job board data against UK visa sponsor registry. Data processing application with practical developer utility.
Fuzzy-matching tool for matching job board data against UK visa sponsor registry. Data processing application with practical developer utility.
Web app for AI-powered video generation with intuitive UI. LLM/generative AI application but limited developer tooling focus.
Collection of prompts and responses from DeepMind's Aletheia model on advanced mathematics problems.
iOS app converting text prompts to complete songs using AI models. LLM application but music-focused, limited developer relevance.
Short article discussing building reusable AI agent skills to avoid repetitive instruction setup across sessions with Claude Code.
GPT-5.2 proposes new gluon amplitude formula in theoretical physics, later formally proved by OpenAI and academic collaborators.
ChatGPT adds Lockdown Mode and Elevated Risk labels to defend against prompt injection and AI-driven data exfiltration.
OpenAI describes real-time access system combining rate limits and credits to power continuous access to Sora and Codex.
GABRIEL is an open-source toolkit using GPT to convert qualitative text and images into quantitative data for social science research.
MedXIAOHE: medical vision-language foundation model with entity-aware pretraining for clinical applications, achieves SOTA on medical benchmarks.
Research comparing small language models vs large language models, focusing on task-optimized, on-device variants with reduced computational requirements.
Xiaomi-Robotics-0: open-sourced vision-language-action model for robotics with cross-embodiment pretraining and optimized real-time execution.
Seedance: Next.js-based tool generating videos and images from text prompts using AI with template and credit-based pricing.
RL-based sim-real co-training method for vision-language-action models that combines simulation and real-world robot data using reinforcement learning instead of supervised fine-tuning.
Codeman: open source security-focused launcher for Codex with permission levels, session resumption, and webhook notifications.
User study examining explainable AI methods in no-code ML platforms, addressing transparency gaps for non-technical users in sensitive domains.
Latent Generative Solvers framework using VAE and Transformer for long-horizon physics simulation across heterogeneous PDE systems with uncertainty mechanisms.
Reference architecture for securing enterprise AI estates with multi-agent governance covering LLMs, RAG pipelines, agents, and shared infrastructure security.
Bi-level prompt optimization method aligning LLM-based evaluations with human judgments without costly fine-tuning, applicable across tasks and datasets.
AgentNoiseBench benchmarks robustness of tool-using LLM agents under realistic noisy conditions versus idealized settings to identify deployment gaps.
Agentic RL approach optimizing proactive LLM agents' Pareto frontiers through behavioral optimization for efficient multi-turn task completion.
ReplicatorBench evaluates LLM agents on replicability of social/behavioral science papers, addressing data availability variations and reproduction vs replication gaps.
Object-centric world model extending masked joint embedding prediction with object-level latent interventions for capturing interaction-dependent dynamics.
TRACER framework for estimating uncertainty in multi-turn agentic interactions by detecting trajectory-level breakdowns like looping and tool misuse.
Value-factorization method for cooperative multi-agent RL addressing sim-to-real gaps and environmental uncertainties while maintaining decentralized execution.
Analysis of visual-textual coupling in multimodal RL for MLLMs showing only ~15% of tokens exhibit strong cross-modal connectivity during reasoning.
First full-stack benchmark measuring privacy leakage in multi-agent LLM systems across inter-agent messages, shared memory, and tool arguments covering 1,000 scenarios.
Framework for continuous learning of internal reasoning processes and action scheduling policies in AI systems for sustained adaptation in dynamic environments.
Multi-agent conversational system automating causal inference workflows by lowering technical barriers and handling algorithm selection, data quality, and result interpretation.
Framework for tool-augmented LLM agents operating under monetary budget constraints using intention-based planning to handle expensive tool invocations.
Method to automatically configure LLM-based agents by learning query-wise decisions for workflows, tools, token budgets, and prompts instead of using fixed templates.
Survey of multi-agent communication in MARL, robotics, and LLM systems covering who communicates, what/when/why communication occurs, and emergent language.
MAPLE improves RL post-training for multimodal language models by weighting modalities based on task relevance, reducing variance and improving robustness to missing/added signals.
scPilot framework enables LLMs to perform omics-native reasoning on single-cell RNA-seq data, automating cell-type annotation and trajectory reconstruction through tool integration.
Studies behavioral consistency of LLM-based agents across 3,000 runs, finding 2.0-4.2 distinct action sequences per 10 runs and showing variance predicts task failure.
Evaluates multimodal LLMs on mathematical spatial reasoning tasks, finding significant gaps (60% vs 95% human accuracy) in 2D/3D relation parsing and manipulation.
Multi-dimensional alignment framework for medical LLMs addressing RLHF limitations in high-stakes applications through verifiable rewards and collaborative optimization.
PhyNiKCE combines neurosymbolic reasoning with agentic LLMs for computational fluid dynamics, addressing probabilistic LLM limitations through symbolic constraints and physics-aware planning.
Benchmark Health Index framework audits LLM evaluation sets to detect score inflation and selective reporting, providing systematic quality assessment of benchmark reliability.
Identifies causal origin of LLM shortcuts through Rung Collapse theory, where autoregressive training fails to distinguish association from intervention, proposing epistemic regret minimization.
Vector-to-Graph pipeline converts CAD diagrams to structured graphs to improve multimodal LLM understanding of engineering schematics and topology, addressing pixel-driven limitations.
ThinkRouter improves reasoning efficiency by routing thinking between latent and discrete reasoning spaces based on model confidence dynamics analysis.
Sparse complementary fusion method for distribution-aware LLM merging that reduces parameter interference and improves generalization over existing weight-space heuristics.
Crosscoders enable unsupervised cross-architecture model diffing to identify differences in LLM internal representations and uncover safety-critical behaviors.
Text2GQL-Bench benchmark for evaluating LLM systems translating natural language into graph query language for graph database analysis and manipulation.
AIR: First incident response framework for LLM agent systems defining protocols for responding to, containing, and recovering from agent failures after deployment.
TSR (Trajectory-Search Rollouts) training approach for multi-turn RL of LLM agents that improves exploration and avoids mode collapse with sparse/delayed rewards.
RL-enhanced LLM framework for advertising text generation that jointly optimizes for both generation quality and online performance metrics like CTR.
FlowMind framework translates LLM reasoning and tool use into structured workflows using execute-summarize mechanism to avoid interference between task solving and workflow generation.
LAVES: LLM-based multi-agent system for generating instructional videos from educational problems with hierarchical decomposition and strict logical rigor.