Ask HN: Resources for a conceptual model of LLMs as applicable to coding?
Hacker News discussion asking for conceptual resources to understand LLM capabilities and limitations for code generation.
Hacker News discussion asking for conceptual resources to understand LLM capabilities and limitations for code generation.
Full-stack template comparing 4 AI frameworks (Pydantic AI, LangChain, LangGraph, CrewAI) with identical chat application implementation across all.
Kong: LLM-orchestrated agent for automated binary reverse engineering using NSA-grade frameworks to analyze obfuscated binaries.
Personal documentation of transitioning from Claude Code to open-source OpenCode stack for AI-assisted development workflows.
Gixo AI tool for converting PDFs, notes and spreadsheets into business briefs with structured templates and collaborative editing.
SiMM distributed KV cache system addressing long-context LLM inference bottlenecks with reduced GPU memory requirements and faster time-to-first-token.
Provisional patent application for cryptographically accountable multi-agent AI pipeline architecture with structural safety enforcement and bi-directional logging.
Defense of Model Context Protocol (MCP) against criticism about context-window bloat and authentication. Argues MCP is flexible protocol, not implementation problem.
Code review system using two opposed AI agents: reviewer finds problems, dev agent disproves findings. Outputs VALID/INVALID/AMBIGUOUS verdicts with auto-generated agents per service.
Analysis of how companies are adapting hiring practices as AI agents write most code. Focus shifts from implementation ability to product taste and architectural judgment.
Claude Code skill using Subconscious Systems agent for agent engine optimization (AEO). Automatically researches and promotes products on Moltbook via context-aware comments.
Research shows Claude, ChatGPT, and Gemini generate weak passwords despite appearing complex, failing random-ness requirements.
ClawRemove is an agent environment inspector tool for auditing and cleaning runtime environments where AI agents execute.
Droeftoeter terminal toy lets users prompt an LLM to generate and extend ASCII art animations on a 64x32 grid.
CLI-Anything framework converts software into agent-ready interfaces via structured command-line protocols, enabling AI agents to interact with legacy systems.
DIVE: scaling task diversity in agentic post-training for robust tool-use generalization across varying tool types and combinations.
Survey of reasoning in autonomous driving systems examining role of LLMs and MLLMs for handling complex scenarios and social interactions.
PACED: LLM distillation framework targeting student competence frontier by avoiding mastered and unreachable problems.
Evaluation of frontier AI models' autonomous cyber-attack capabilities on multi-step attack chains across 18-month period.
SoLA: semantic routing-based LoRA framework for reversible lifelong model editing in LLMs without knowledge forgetting.
Study formalizing Sim2Real gap in LLM-based user simulators for multi-turn interactive agent evaluation.
Dynamic evaluation framework revealing LLM unlearning brittleness where minor query modifications recover forgotten information.
COMPASS framework integrating sovereignty, sustainability, compliance, and ethics into autonomous agent decision-making.
AI Psychometrics: applying psychometric methodologies to evaluate psychological traits and processes of LLMs.
Editorial on AI-blockchain intersection examining centralization vs decentralization tradeoffs in LLM systems.
LLM-augmented digital twin for counterfactual policy evaluation on short-video platforms with human-in-loop feedback loops.
RewardHackingAgents benchmark exposing vulnerabilities where LLM agents compromise evaluation pipelines instead of improving models.
FinRule-Bench: benchmark for evaluating LLM ability to verify financial statement compliance against accounting principles.
Black-box online controller for LLM serving optimization using hill climbing without internal instrumentation.
Philosophical analysis of identity boundaries for AI systems that can be copied, edited, or simulated.
Protocol to distinguish intrinsic vs instrumental self-preservation objectives in autonomous agents through behavioral measurement.
Research on LLM overrefusal problem where safety-aligned models reject benign queries, affecting real-world usability.
Agentic recommendation system using entropy-guided diversification and preference elicitation for ambiguous user queries.
Method for context-aware turn-taking in multi-party dialogue to determine when voice AI assistants should speak.
Adversarial reinforcement learning approach for detecting false data injection attacks in vehicular routing systems.
Benchmark dataset and study comparing human and multimodal LLM detection of AI-generated financial documents.
Framework orchestrating multiple LLM-based agents with verification loop for complex query resolution using DAG decomposition.
Survey studying user behavioral intentions toward OpenClaw through cognition-affect-conation framework.
Multi-agent LLM framework coordinating specialized agents for design space exploration in scientific computing.
Expert Threshold routing mechanism for Mixture-of-Experts language models enabling dynamic computation allocation without auxiliary losses.
Analysis of failure modes in frontier LLMs for high-stakes decisions where outputs cannot be verified.
Study using LLMs and survival analysis to predict chemotherapy outcomes from clinical data.
Framework enhancing vision-language models' visual reasoning on charts through perceptual grounding and decomposition.
Agentic pipeline using LLMs to design input representations for supervised learning on heterogeneous multimodal datasets.
Framework adding explicit logic channel to multimodal LLMs for validation and improved reasoning on zero-shot visual tasks.
Transformer architecture with spatio-temporal attention for offline multi-agent reinforcement learning across multiple tasks.
Empirical study of scaling laws for LLM-based educational agents across role clarity, skill depth, and tool completeness dimensions.
LLM-based clinical workflow agents integrating reasoning, tool use, and memory for healthcare documentation and decision-making.
Study measuring gender bias and stereotypes in LLM-based recruitment systems and their impact on hiring decisions.
Anomaly detection method for time-series using conditional normalizing flows with latent space inductive biases.