NASimJax: A GPU-Accelerated Policy Learning Framework for Penetration Testing
NASimJax: GPU-accelerated RL framework for training penetration testing agents with realistic network simulation.
NASimJax: GPU-accelerated RL framework for training penetration testing agents with realistic network simulation.
Learning framework for evolving models through user interaction queries; theoretical foundations for deployed systems.
TransXion benchmark for anti-money laundering detection using realistic transaction graph datasets.
Deep reinforcement learning for CO2 storage control with latent model adaptation under partial observability.
Sutra: compiler for vector symbolic architectures that targets PyTorch neural networks with tensor operation fusion.
Analytical solution to Mountain Car problem revealing optimal control simplicity and introducing Chebyshev policies as universal RL policy class.
Theoretical analysis of mini-batch scaling laws in sketched linear regression across single-pass and multi-pass SGD settings.
Benchmark evaluating LLM zero- and few-shot performance on binary tabular classification without labeled context examples.
Deep reinforcement learning overlay for pair trading strategy in cryptocurrency markets with high volatility adaptation.
Benchmark for evaluating AI agent capabilities across diverse environments beyond common applications, addressing limitations of saturated performance on existing benchmarks.
Parameter-efficient adapter approach for knowledge editing in LLMs using memory retrieval and dual routing mechanisms to update facts while preserving model behavior.
Signature filtering module enhances statistical watermark detection in LLM outputs without modifying generation or embedding.
Generalization Spectrum framework evaluates learning algorithms on per-sample transfer ability rather than aggregate test scores.
MiniOpt framework enabling LLMs to reason, model and solve diverse optimization problems with minimal training resources.
Study of adversarial robustness of AI-generated image detectors, testing methods against evasion and poisoning attacks.
Reconstruction Alignment (RECA) method improves unified multimodal models by leveraging visual understanding for better generation.
Self-supervised contrastive learning method for patent document representation, optimizing dropout and temperature settings.
Semi-supervised few-shot learning approach using vision-language models and auto-annotation for learning from limited labeled data.
Study showing metaphors in training data cause cross-domain misalignment and reasoning errors in large language models.
DEFault++ hierarchical fault detection system for identifying component-level failures in transformer architectures without visible errors.
Weak model drafts improve strong model GRPO training by injecting mathematically wrong domain-specific examples, outperforming standard on-policy RL.
Benchmark evaluating deep research agents on expert consulting tasks with 70 SME-authored prompts embedding cognitive traps.
Multi-agent LLM systems with symbolic reasoning frameworks show emergent risk behavior changes through memory interactions and reflective prompts.
arXiv paper developing framework for analyzing representation costs in parametric data-fitting and deep neural networks via regularizers.
arXiv paper on agentic LLM workflows for zero-shot information extraction from lung pathology reports using prompt-plan-extract approach.
arXiv paper on A-Evolve-Training, autonomous post-training system for 30B model with no human-in-loop over multiple weeks.
arXiv paper on GRAG framework for personalized conversational agents with separate treatment of content grounding and personalization.
arXiv paper evaluating reliability of automated jailbreak judges for LLM safety, comparing classifiers vs human evaluation.
arXiv paper introducing Autodata, an AI agent that acts as data scientist to generate synthetic training data via Agentic Self-Instruct.
Technical guide on building domain-specific expert chatbots using Claude with architectural patterns and two-model approach.
GPU/VRAM filter tool for searching which LLMs will run on specific hardware at various quantization levels.
Discussion of how ML engineering roles have evolved as frontier LLMs replace custom-trained models in non-lab environments.
Terminal UI tool for building and testing shell pipelines interactively with live output and diff capabilities.
White House requests OpenAI restrict GPT 5.6 release to government partners due to advanced capabilities.
Discusses guardrails and safety mechanisms for preventing AI agents from executing harmful instructions.
Local-first AI coding assistant for JetBrains IDEs with offline inference capabilities.
Ludion is a routing system for AI inference that optimizes WebGPU behavior and resource allocation.
Anthropic accuses Alibaba of orchestrating largest AI model distillation attack with 28.8M fraudulent exchanges.
DropItDown: macOS tool converting files to clean Markdown locally for AI agents without account or token burn.
Extropic announces breakthrough AI algorithms and hardware addressing energy efficiency as limiting factor for AI scaling.
CS2-10k dataset: large-scale egocentric Counter-Strike 2 video with synchronized action inputs for world model training.
HoprLab: Python CLI toolkit for simulating AI training math, estimating model size, VRAM usage, training time, and config risks before actual model training.
DeepSeek Flash inverted AI agent economics: open-source model beats Sonnet on benchmarks, eliminates pricing subsidy dynamics favoring developers.
macOS malware 'Gaslight' embeds fake errors and prompt injection strings to confuse AI-assisted malware analysis tools.
Chrome extension for exporting Claude.ai conversations, artifacts, and visible thinking as organized ZIP files with local browser processing.
JetSpec: LLM inference optimization achieving up to 9.64x speedup and 1000 tokens-per-second throughput using speculative decoding.
Study comparing general LLMs versus specialized clinical AI tools on medical benchmarks. No detailed findings or data provided.
Telnyx AI Inference service extracts structured JSON from unstructured text like support tickets and emails using configurable schemas.
Technical analysis of scaling laws in deep learning, their empirical foundations, and optimal compute allocation strategies.
Discussion: Users sharing experiences giving AI agents phone numbers and real-world tool access for task automation.