Bound to Disagree: Generalization Bounds via Certifiable Surrogates
Theoretical work on generalization bounds for deep learning using disagreement-based certificates and surrogate models.
Theoretical work on generalization bounds for deep learning using disagreement-based certificates and surrogate models.
Decentralized federated learning framework addressing data heterogeneity and convergence challenges in server-free collaborative learning.
Novel attribution method (DPA) for understanding internal mechanisms of transformer LLMs, balancing faithfulness and computational efficiency.
Test-time adaptation method for LLMs addressing distribution shift through synapse consolidation inspired by biological pathways.
ERA framework for embodied agents using event-centric world modeling and memory-augmented retrieval for interpretable decision-making.
Analysis of solver vs. sampler roles in LLM-based simulations, showing models optimally solve rather than realistically sample behavior.
Neural network approach to approximate attention computations over cached contexts, reducing per-token computation cost.
Trust-region fine-tuning method addressing compounding occupancy shift failure in sequential multi-agent LLM team updates.
Parallel inference method for transformers using Newton corrections to relax sequential layer dependencies and reduce latency.
Studies substitution of human-curated tasks with synthetic augmentation for reinforcement learning training of agentic language models.
Empirical study showing LLMs fail to evaluate source quality during multi-source synthesis despite isolated fact-checking capability.
Framework for stabilizing multi-agent LLM coordination through entropy-regularized equilibrium selection in incomplete-information games.
Method for detecting and steering concepts in transformer models using raw hidden state dimensions without training.
Weak-to-strong generalization via on-policy distillation: train RL on smaller models then transfer to scale stronger models efficiently.
In-context learning paradigm for system identification enabling one-step prediction and multi-step simulation from class-level observations.
Fast statistical method for localizing watermarked segments in LLM-generated text using epidemic change-point detection.
LiveOIBench: Large-scale competitive programming benchmark evaluating LLM coding capabilities with challenging problems and comprehensive test coverage.
Tutorial review of diffusion models for simulation-based inference with applications to parameter estimation from simulated and real data.
LLM-informed model-based planning framework for object search in partially-known environments using prompt selection methods.
Research on amortized inference for discrete choice models using equivariant neural networks, addressing logit model limitations.
VTC compiler eliminates data movement in DNN execution via virtual tensor optimization, applicable to LLMs.
Analysis of visual counting failures in Vision-Language Models, identifying bottlenecks in systematic generalization.
TLA-Prover uses 20B parameter LLM with LoRA fine-tuning to synthesize formally-verifiable TLA+ specifications.
RhinoVLA optimizes Vision-Language-Action models for real-time robotic deployment on edge hardware.
DYNA-PRUNER framework for input-adaptive model pruning in spatio-temporal prediction tasks.
MTEB-BR benchmark for evaluating text embedding models on 22 Brazilian Portuguese tasks.
Academic tool personalizes AI-generated research papers and grant proposals by erasing signs of AI authorship and tailoring tone.
Open source project running Llama 2 LLM on vintage DOS machines from 1996-2004 with modern hardware variants using FreeDOS.
Guide to Character.ai alternatives offering better personality, memory, and creative freedom for roleplay conversations.
Interactive atlas mapping 8.5M research papers with LLM summaries, citations, and peer reviews using UMAP embedding and WebGL visualization.
Tool that compiles recorded LLM-agent behavior into verified WebAssembly binaries for reproducible agent execution.
Developer shares experience of LLM usage fatigue despite moderate adoption; discusses workflow with Claude and Codex for coding tasks.
Analysis showing public LLM benchmarks are unreliable; Claude Opus/Sonnet performance varies significantly across real workloads compared to rankings.
KPMG survey finds 30% of corporate leaders struggle understanding AI implementation costs as vendors shift to usage-based pricing models.
Zoom app with real-time AI guidance for sales calls, helping with discovery questions, objections, and sales playbook alignment.
Research or analysis on security vulnerabilities in defensive AI agents, showing remote code execution attack vectors.
LLM-generated blog where Fable writes articles, code, and SVG from its own POV, demonstrating creative capabilities.
vLLM transformer inference backend optimization for native-speed performance.
Obsidian vault with remote sync, MCP support, and API. Developer tool for note-taking.
Actenon: open-source authority broker for AI agents with capability scoping, approval gates, and runtime limits.
Runtime security and capability scoping for AI agents. Research from Harvard/CMU identifying sandbox escape vulnerabilities.
Non-custodial spot market for AI inference tokens, designed for agent-to-agent transactions on-chain.
Dashboard to manage multiple Macs as Claude Code agent fleet. AI agent orchestration tool.
Istota: self-hosted personal AI operating system with multi-agent capabilities, cloud integration, and modular skills.
Prompt compiler with reasoning transparency and governance for controlled LLM execution.
Tool to convert document collections into searchable knowledge base. LLM application for document processing.
Nully: Minimal open-source AI chat interface in Go with simple messaging and history, no advanced features.
Research notes on agentic coding processes, test automation approaches, and LLM benchmarks (insufficient content provided).
Onboard-CLI: AST-based command-line tool using LLMs to visualize and map complex codebases via Tree-sitter parsing and React Flow canvas.
Skill-extractor tool extracts reusable skills from coding agent transcripts for automated task automation.