BioArc: Discovering Optimal Neural Architectures for Biological Foundation Models
Neural architecture search for biological foundation models considering physicochemical properties unique to biology.
Neural architecture search for biological foundation models considering physicochemical properties unique to biology.
Sequential hypothesis testing framework for reliable verification of AI agent trajectories and action quality scoring.
Analyzes asymmetric realizability problem when transplanting tokenizers across LLMs, studying coefficient reconstruction over shared lexical anchors.
Uses LLMs to recover semantic relationships in categorical data clustering by overcoming limitations of co-occurrence statistics.
Circuit analysis comparing computational mechanisms of autoregressive models versus masked diffusion language models post-trained from the same base.
Noise-compensated sharpness-aware minimization optimizer for learning from noisy labels by connecting label noise to SAM flatness-seeking behavior.
Proposes entropy-based metric (HE-SNR) for guiding mid-training of LLMs on software engineering tasks, addressing limitations of perplexity metrics.
Introduces Representation Unlearning framework for machine unlearning via information compression in representation space rather than parameter modification.
Analyzes block Hadamard rotations in post-training quantization for neural networks, providing systematic non-asymptotic analysis of outlier suppression.
Framework for extracting interpretable rationales from DNNs via knowledge distillation using select-predict architecture with remote supervision.
Shows that SFT optimization for downstream RL performance differs from isolated SFT performance, revealing checkpoint selection mismatch in reasoning LLM post-training.
Proposes Rectified Distribution Matching Regularization for Joint-Embedding Predictive Architectures to learn sparse representations while preventing collapse.
Research on how LLMs encode future reasoning in hidden states versus explicit chain-of-thought steps, exploring the latent planning horizon in multi-step reasoning.
FedNMap method for composite federated learning with nonsmooth regularizers, achieving linear speedup across heterogeneous distributed systems.
Learning-based approach to aerodynamic inverse design balancing performance optimization with design language preservation.
Discrete diffusion sampling algorithms for off-policy sampling from unnormalized densities, with applications in latent space modeling.
Framework for systematic analysis of rare events and edge-case behaviors in large language models during deployment, addressing gaps between development and production.
Gradient preconditioning method for efficient reward-guided generation with one-step models, addressing reward hacking and speed issues in test-time optimization.
Claude Code Dynamic Workflows feature enables resumable multi-day parallel tasks with preserved progress, potentially replacing manual context management.
Gitea security vulnerability CVE-2026-27771 allowed unauthenticated private container image access affecting 30k+ deployments for 4 years.
Unknown LLM model 'Hy3' ranking high on OpenRouter leaderboard.
Recommendations for designing independent third-party evaluations of frontier AI model capabilities and safety mitigations.
Research on building reliable LLM-based judges for evaluation tasks.
Brinicle: C++ vector database engine with Python wrapper achieving sub-millisecond latency on 1.2M items, supporting hybrid semantic/lexical search.
Local-first desktop app using semantic search to find files by plain English queries instead of exact filenames, built with ML.
Protocol for multi-agent LLM deliberation systems with replay capabilities for agent coordination.
Tool to convert code folders into single LLM prompts for easier AI processing and context.
User reports Claude Opus 3.5 performance degradation with file reading errors and hallucinated commands in typical development workflows.
User discussion on whether Claude Code reduces need for frontend frameworks like React in building complex web applications.
Open-source MCP server connecting AI clients to Obsidian vaults for semantic search and writing.
Local-first semantic code graph for AI coding agents via MCP protocol. Runs locally without external APIs.
Open-source AI browser agent that automates repetitive web tasks 24/7 by learning from user actions.
LLM-based journal application that analyzes natural language entries to detect bipolar mood escalation patterns.
Analysis of LLM-generated text patterns showing identical sentence structures spreading across internet, identifying systemic AI writing artifacts.
Open-source tool (aw team bootstrap) for structuring multi-agent teams with global IDs, communication, roles, and automation from templates.
Enough is extensible personal language system with local models and OpenRouter support for planning, writing, and translation tasks.
Ktx is open-source executable context layer for reliable data agents, addressing accuracy issues in SQL generation and warehouse interactions.
Custom Claude Code agent skill for enforcing conventional commit message format with validation hooks.
Research on enhancing LLM code analysis with formal reasoning engine for transitive code questions and dead code detection.
Analysis of failures in Google's AI Overview demonstrating character counting and spelling errors in simple tasks.
Analysis of enterprise security risks from autonomous AI agents with 66% benchmark accuracy increase, discusses control mechanisms.
Genesis Architect scans 15-20 GitHub repos to extract architecture patterns, bug issues, and security patches, then generates project scaffolds based on production data rather than templates.
Opinion piece on how AI-generated content proliferation made creation easier but quality assessment harder across media formats.
Analysis of AI-generated malware npm package targeting Claude users that leaked its own GitHub token via security misconfiguration.
Familiar: open-source screen/clipboard capture tool using Apple's local OCR, providing context for AI agents via markdown output.
rotom: local Rust gateway enabling Codex and Grok OAuth with OpenAI/Anthropic-compatible APIs, allowing use of alternative providers.
Guide for making Symfony applications agent-ready with markdown negotiation, OpenAPI docs, and API exposure for AI agent integration.
VS Code extension for selecting code blocks by syntax structure rather than manual dragging, with AST-aware navigation.
SEO automation toolkit using NodeHub API and optional LLM for SERP analysis and cluster naming. Promotional package by Senuto.
Endava uses Codex to accelerate requirements analysis and software delivery, reducing analysis time from weeks to hours.