HN OsamaJaber 28d ago

Build You Own Model

RunInfra benchmarks and optimizes LLM model serving. Compares inference engines, GPU allocation, latency, throughput, and provides deployment kits.

HN thunderbong 28d ago

10x smaller vector indexes in pgvector

TurboQuant vector quantization reduces pgvector index size by 10x. Optimization for vector databases used in LLM applications.

HN Gedxx 28d ago

Transcribe.cpp

transcribe.cpp: Open-source C/C++ speech-to-text inference library with GPU support. Portable STT tool similar to llama.cpp.

HN Robert_Linz 29d ago

TS Foundation Models

TiRex-2 pretrained time series foundation model for zero-shot multivariate forecasting with streaming support.

HN kristianpaul 29d ago

Deepagents

Deepagents: Open-source opinionated agent harness library with extensible components, includes preconfigured coding agent for terminal use with any LLM.

HN arberx 29d ago

Every AI Visibility Tool Is Lying to You

Analysis of AI visibility measurement tools showing dashboard metrics for brand mentions in ChatGPT, Claude, Gemini lack methodological support.

Ax Aryuemaan Kumar Chowdhury, Afreen Shaik, Yaparla Bhargavi, Brahma Kumar 29d ago

The Wiola Architecture for Efficient Small Language Models

Wiola is a novel small language model architecture from first principles with spiral rotary positional encoding and gating mechanisms for efficient inference.