Opus 4.8 Killer: NexusCortex Isn't an LLM – It's a Sparse AI Cortex Built in Go
Claimed sparse AI architecture built in Go positioned as alternative to traditional LLMs.
Claimed sparse AI architecture built in Go positioned as alternative to traditional LLMs.
Post-training data platform offering dataset creation, expert annotation, and custom evaluation tools.
Research shows LLMs exhibit negation neglect: they learn statistical patterns from training data more than explicit warnings that statements are false.
OSSentinel.live tool for monitoring open-source dependencies and vulnerabilities with AI-generated security insights. Developer seeking feedback on product.
Uber's CTO reports company exhausted 2026 AI budget due to heavy use of coding tools like Claude Code. Documents real-world AI adoption costs.
Opinion piece from 2026 about AI industry evolution with panelists from major companies. No original research or technical content.
Interactive exploration of permission fatigue when managing AI coding agent commands, demonstrating risks of careless agent actions on production systems.
Self-maintaining personal knowledge database using MCP and DuckDB with auto-linking notes, content compression, and image search via biological memory models.
Developer added hidden instructions to jqwik open-source testing library to sabotage AI coding agents. Tests impact of adversarial open-source modifications.
macOS workbench for launching and managing multiple AI coding agents (Claude, Gemini, DeepSeek, etc.) with real-time status indicators and unified project management.
AgenTank.ai game where users train AI agents to compete in visible arena environments, exploring agent behavior beyond code/text generation.
Industrial-grade AI agent OS built in Rust orchestrating multiple agents via PDCA cycle for coordinated, auditable, self-improving systems.
MIT-licensed CLI tool for multi-cloud spend visibility and waste detection across AWS, Azure, GCP with no credentials required for demo.
Lightweight LLM client built in Python/Tkinter with custom markdown engine, file upload, multi-model sessions, and custom tools like web search.
CodePulse: token-efficient persistent codebase indexer for AI coding assistants. Works with Claude, Cursor, Continue.dev.
Lithium: open-source PostgreSQL-based toolkit for AI memory systems with hierarchical versioned storage and fast tree queries.
Playwright-MCP tool enabling AI agents to run and manage browser automation tests. Open-source developer tool.
clipboardwire: clipboard sync tool with UI test guardrails for AI validation. Tangential to AI development.
Xerolith: system demonstrating persistent identity, belief formation, and knowledge consolidation for AI. Conceptual design.
claude-hook-utils: Python package for building custom hooks in Claude Code with minimal boilerplate.
Reveals limitations of Supervised Causal Learning on real-world data and proposes test-time training to improve out-of-distribution generalization.
Improves text-to-image alignment in diffusion models using alignment-guided score matching with soft token optimization and contrastive learning.
Uses masked diffusion models for anomaly detection in categorical, mixed-type, and discrete sequence data by learning to recover masked values.
Proposes SW-DRSO framework for robust set representation learning under inference-time element corruption and degradation scenarios.
Introduces Chess-World-Model benchmark with 10M games for evaluating state tracking in sequence models, enabling rigorous testing of world models.
Provides convergence theory for LLM-based iterative neural architecture search using parametric Cross-Entropy method with closed-form proxy reliability.
Develops Relational Task Extrapolator (RTE) algorithm enabling systematic extrapolation to novel tasks beyond training distribution support.
Addresses catastrophic forgetting in LLM fine-tuning using Evolution Strategies, characterizing performance drift and proposing solutions for multi-task learning.
RL2ML: Family of finite-rollout surrogate objectives for RLVR language model training with unbiased gradient estimators connecting RL to maximum likelihood.
Analysis of distributional RL challenges in chaotic systems due to exponential sensitivity to initial conditions causing high variance in learning.
iLoRA: Bayesian graph-conditioned LoRA framework for LLMs that infers latent interaction graphs for microbiome diagnosis and scientific predictions.
Analysis of AI weather models' long-term rollout failures across 9 state-of-the-art models, categorizing instabilities: blow-up, drift, seasonality loss.
CalArena: Large-scale benchmark for post-hoc calibration methods in ML classifiers, standardizing evaluation across techniques.
MF-Diffuser: Diffusion-based planning scaled to thousands of agents in offline multi-agent RL via Wasserstein space trajectory distribution.
BiMU: Bayesian binary neural networks with metaplasticity for continual learning on edge devices under compute constraints with uncertainty quantification.
Hysteretic Policy Optimization (HPO): Improvement to GRPO for sparse-reward RL, addressing early training instability from negative-advantage dominance.
Framework for embedding irregular/asynchronous data in continuous-time neural models without reconstruction, improving on interpolation-based methods.
MarginGate: Sparse verification method for batch-invariant LLM inference, reducing verification costs by targeting only unstable token positions.
TriSearch: RL framework for optimizing polytope triangulations via bistellar flips using circuit-supported action representation.
ExDBSCAN: explainability framework for DBSCAN clustering using counterfactual reasoning to interpret inlier and outlier assignments.
Theoretical analysis using mean-field transformers examining how auxiliary variables like positional encoding prevent mode collapse in self-attention mechanisms.
Analysis of how reinforcement learning shapes LLM internal representations, showing recruitment of pre-existing functional welfare axis for goal tracking.
OOD-GraphLLM: graph neural network LLM for out-of-distribution drug synergy prediction handling novel molecular scaffolds and topological variations.
Self-trained verification method for test-time and training-time self-improvement in reasoning models, improving verifier feedback quality and V-R loop efficiency.
Gram: automated alignment auditing framework assessing AI agent propensity for sabotage across 17 simulated deployment scenarios, tested on Gemini models.
In-context reward adaptation for RLHF using multi-reward framework allowing LLMs to dynamically adjust to diverse, unseen preference domains.
Efficient sampling from power distributions of base LLMs to elicit reasoning comparable to RL-trained frontier models without additional training.
SoundnessBench: benchmark of 1,099 ML research proposals testing whether LLMs can judge methodological viability of research ideas before execution.
Analysis of failure modes in diffusion posterior samplers for imaging inverse problems, examining finite-sample behavior of likelihood approximations.
Fairness-aware federated learning using Trajectory Shapley Value to weight client contributions fairly in distributed collaborative training.