Revisiting the LiRA Membership Inference Attack Under Realistic Assumptions
Re-evaluation of LiRA membership inference attacks under realistic assumptions, questioning prior effectiveness claims with realistic threat models.
Re-evaluation of LiRA membership inference attacks under realistic assumptions, questioning prior effectiveness claims with realistic threat models.
Systematic comparison of four training objectives (cross-entropy, prototype, triplet, AP loss) for out-of-distribution detection in image classification.
Security analysis of large vision-language models vulnerable to semantic slot filling attacks that elicit unsafe outputs.
Hierarchical multi-agent system for Kubernetes autoscaling addressing resource waste through coordinated pod and node scaling policies.
arXiv paper on Staged Multi-Agent Training (SMAT) for co-adaptive exoskeleton control, using curriculum learning to mirror human motor adaptation.
arXiv paper on silicon photonics acceleration for diffusion model inference, targeting energy efficiency of UNet and attention mechanisms.
arXiv paper on physics-based reinforcement learning for data-driven exoskeleton control using joint-moment prediction instead of lab-based inverse dynamics.
arXiv research evaluating synthetic data for baggage trolley detection in airport logistics systems.
arXiv research on federated learning with compression for non-convex optimization on heterogeneous distributed data.
arXiv research on ML-driven microarchitectural techniques addressing memory bottleneck in modern computing systems.
arXiv paper on scaling Mixture-of-Experts model training using Megatron Core, addressing systems challenges in sparse model architectures across memory, communication, and computation.
Registry of AI prompt configuration files (.md format) for sharing and discovering real LLM setups and workflows.
OAuth 2.1 auth provider for FastMCP enabling personal MCP servers on Claude web, mobile, desktop, and Code without external identity providers.
Context Hub provides curated, versioned markdown documentation for coding agents to reduce hallucinations and improve API knowledge retention across sessions.
Deterministic safety proxy for MCP servers without using LLMs. Blocks unsafe tool access before Claude/Cursor invokes them.
TypeScript MCP server starter kit with auth, rate limiting, AWS CDK, Docker. Enables building custom AI tools for Claude/Cursor in minutes.
PyTorch compatibility layer redirecting CUDA calls to AMD/Intel/Ascend backends. Runs ML workloads on non-NVIDIA hardware without code rewrites.
Report on Chinese AI companies distributing 8 billion yuan in coupons during Lunar New Year for agentic AI apps. Market analysis of agent deployment in China.
Postmortem of Tess.Design, AI image marketplace with artist royalties (50%). Launched May 2024, shut down January 2026 with learnings on ethical AI models.
Study evaluating 14 AI agents across 2 benchmarks on 12 metrics across 4 reliability dimensions. Finds recent capability gains yield only small improvements in actual reliability compared to accuracy scores.
Open source agent framework forking OpenAI's Symphony, using Claude Code for autonomous implementation of Linear board issues. AI agents with LLM integration.
Open source model-agnostic AI code review tool with full control over model choice and costs. Alternative to Claude Code Review.
Andrej Karpathy thought piece on autonomous AI agents conducting frontier research across compute clusters. Speculative/fictional framing of agentic research systems.
Security research on model artifact integrity during local LLM inference in llama.cpp. Creates llm-inference-tampering project targeting inference-layer attacks.
Framework for defining requirements and specifications for AI systems beyond testing/evals. Addresses gap between eval scores and actual user satisfaction in AI products.
PUG: tool that converts messy API documentation into structured CLI tools and MCP servers using LLMs for AI agents.
TLAi+ Benchmarks: dataset and benchmark suite for evaluating LLMs on TLA+ formal specification tasks with diverse problem types.
Autonoma: AI agents that automatically generate test suites and find bugs by navigating applications without manual test scripts.
Rainy Updates: deterministic dependency review and upgrade tool for Node monorepos with CI/CD integration and automated fix PRs.
Essay on limitations of open weights models without open training data, discussing post-training challenges for trillion parameter models.
AI-powered technical interview prep tool simulating realistic interviewer interactions with WebRTC and Socket.io.
Nvidia planning to launch NemoClaw, an open-source AI agent platform for enterprise software companies to dispatch agents.
Plannotator: open source tool for manual code review and feedback loops for autonomous agents. OSS framework for agent improvement via human feedback.
CLI tool using Claude to analyze project codebases and generate customized Claude Code configurations. Integrates with Claude CLI for code-specific setup.
Analysis of SRAM-centric AI accelerators (Cerebras, Groq, d-Matrix) vs GPUs for inference, focusing on near-compute vs far-compute memory tradeoffs.
Agentis: AI-native programming language with LLM as standard library, using binary hashed DAG for version control instead of text files.
Jobbi.app: AI tool that automatically tailors resumes to job descriptions by extracting relevant content from master resume.
Part 2 of observability-driven harnesses for autonomous optimization of systems built with AI agents, focusing on verification loops.
Analysis of Claude Code's /loop feature enabling autonomous agent operation with multiple roles and cadences for AI programming workflows.
Technical article on prompt injection security vulnerabilities in AI coding agents, using supply chain attack case studies from 2026.
Using autonomous AI agents to automate management of open source repositories. Minimal content provided.
Architecture and benchmarks for context plane infrastructure that provides AI agents with targeted context retrieval via S3, replacing prompt stuffing and multiple API calls.
Research study finds AI-based review monitoring systems reduce angry employee responses to negative customer feedback.
Open-source DAW plugin using JUCE and React; bridges generative music models with professional audio software via Magenta integration.
Microsoft launches Copilot Cowork AI agent for autonomous office work (meeting prep, calendar management, document creation) across Microsoft 365.
Commentary arguing code generation is not the bottleneck in development; contextual analysis of AI productivity gains in practice.
Multiplayer cloud desktop with AI agent sandboxing capabilities; encrypted filesystem allows secure testing of autonomous agents.
Self-hosted OpenAI-compatible LLM gateway with automatic failover, routing, and quota management across multiple providers without code changes.
Lurk is a local agent that provides context to Claude Code, Cursor, and ChatGPT by tracking user activity, eliminating the need to re-explain work context in each conversation.
AI-native healthcare information system with policy-gated clinical agents (triage, orders, lab review, pharmacy). Uses VERITAS trust layer, OPA Rego policies, FHIR R4, and cryptographic audit.