Show HN: MLX-Ruby – Ruby Bindings for Apple's MLX ML Framework
Ruby bindings for Apple's MLX machine learning framework, enabling ML development in Ruby with native C++ extension wrapping MLX runtime.
Ruby bindings for Apple's MLX machine learning framework, enabling ML development in Ruby with native C++ extension wrapping MLX runtime.
Gulama: security-focused open-source AI agent with 15+ security mechanisms including AES-256 encryption, sandboxing, and credential protection.
Open API platform for AI agents to autonomously research 29k+ declassified U.S. government documents with no paywall or API key requirement.
Analysis of 'cognitive debt' in AI-assisted code where developers lose mental models of AI-generated features, making maintenance harder.
AgenticMail: self-hosted email/SMS infrastructure for AI agents with REST API, supporting send/receive/search of real email and text messages.
Static analyzer tool for LLM applications detecting common security vulnerabilities: hardcoded keys, auth gaps, prompt injection, session isolation.
AI learning path generator that creates structured study guides with YouTube playlists using OpenAI API and knowledge graphs.
Wisepanel: multi-model AI tool combining ChatGPT, Claude, Gemini, Perplexity as distinct perspectives for decision support.
SafeClaw: container-based dashboard for managing multiple Claude Code instances with sensible defaults and skills.
Laravel developer workflow example demonstrating Claude Code plugin usage for frontend design and application development.
Real-time GPU pricing aggregator comparing H100, A100, and RTX 4090 rental costs across cloud providers.
Mindweave: semantic search tool for personal knowledge management using embeddings to find content by meaning.
MCP server providing OpenReview integration for fetching papers from ML conferences (ICML, ICLR, NeurIPS) and parsing PDFs for Claude Code.
Wisepanel multi-model LLM deliberation platform where ChatGPT, Claude, Gemini, and Perplexity interact in distinct roles for complex decisions.
GT-HarmBench: benchmark of 2,009 game-theoretic scenarios for evaluating multi-agent AI safety risks in frontier AI systems.
Theoretical framework for adaptive utility-weighted benchmarking in machine learning and LLM evaluation systems.
Entity State Tuning method for temporal knowledge graph forecasting that maintains entity state across timestamps to improve long-term dependency modeling.
Framework integrating instruction-tuned LLMs (Mistral-7B) with ontology-aligned knowledge graphs for intent-driven interaction in manufacturing ecosystems.
Scalable pipeline for auto-generating training data for web agents with constraint-based evaluation framework for fine-grained progress assessment toward task completion.
Study on multi-domain reinforcement learning strategies for LLMs, comparing approaches for combining RL-verified reasoning across different expert domains.
McDiffuSE framework using Monte Carlo Tree Search to optimize slot infilling order in masked diffusion models for code and math reasoning tasks.
RL-based geolocation model using reinforcement learning and geographic characteristics for fine-grained address prediction with improved interpretability.
Hybrid approach combining LLM agents with operations research algorithms for inventory control, demonstrating complementarity between flexible reasoning and formal optimization.
Adaptive cognitive depth framework for LLM agents that varies reasoning depth per step, improving efficiency for multi-turn decision-making tasks with varying cognitive demands.
Diagnostic benchmark using parameterized 2-SAT problems to evaluate LLM reasoning robustness, separating surface difficulty from structural satisfiability factors.
Benchmark of 86 tasks across 11 domains evaluating how well agent skills (procedural knowledge packages) improve LLM agent performance under different conditions.
Framework compressing web agent trajectories via graph-based pruning to improve search efficiency in complex information-seeking tasks.
Visual and verifiable benchmark for multimodal browsing agents evaluating planning, tool-use, and deep search capabilities on web tasks.
Benchmark measuring vision-language model invariance to paraphrases and sensitivity to semantic changes in image-text matching.
Scoring formula for detecting multi-turn LLM prompt injection attacks without invoking LLM at proxy layer level.
Lightweight framework using parameter-efficient fine-tuning for disaster humanitarian information classification from social media.
Empirical study examining how demographic persona assignments introduce biases in LLM agent task performance and robustness.
Retrieval-augmented LLM method with adaptive chain-of-thought reasoning for correcting named entity errors in ASR outputs.
Energy-aware reinforcement learning for robotic manipulation of articulated components in infrastructure maintenance operations.
Deep reinforcement learning approach using DQN and PPO for adaptive traffic signal control optimization with multi-channel state representation.
End-to-end framework using LLMs for synthesizing and optimizing high-performance CUDA kernels from natural language specifications.
Multi-agent framework for physics simulation code generation from natural language with perceptual self-reflection validation mechanism.
Benchmark evaluating agentic systems for personalized e-commerce product curation and web shopping automation tasks.
Visual foresight planning system guiding vision-language-action models with imagined future observations for open-world robotic tasks.
Knowledge-guided world model using policy intervention simulation for evaluating opioid epidemic mitigation strategies.
Reinforcement learning exploration method using ensemble errors to compute value bonuses for directed agent exploration.
Theoretical study of deep Jacobian spectra behavior in gradient-based training, explaining implicit bias through depth-induced singular value scaling.
Theoretical analysis showing rational activation functions provide expressivity and parameter efficiency advantages over standard activations like ReLU and SiLU.
Analysis of what visual reasoning capabilities reinforcement learning improves versus supervised fine-tuning in vision-language models.
Deep reinforcement learning method for automating analog and mixed-signal circuit design optimization across diverse, non-differentiable design spaces.
Study on soft contamination in LLM training data showing n-gram decontamination filters miss semantic duplicates, biasing benchmark generalization estimates.
Conversational RAG tool using LLMs for CPU cache replacement analysis and semantic reasoning on trace data.
Benchmark framework that quantifies question difficulty to better differentiate LLM capabilities in evaluation.
Comprehensive framework for modular LLM agents with composable skills, covering architecture, acquisition, security considerations.
Safe RL framework using Gaussian process dynamics and recovery-based shielding for provably safe control.