GLM-5.2's Code Reviews Are Only as Good as Your Prompt
Analysis of GLM-5.2 open-weight model's inconsistent code review performance and prompt sensitivity.
Analysis of GLM-5.2 open-weight model's inconsistent code review performance and prompt sensitivity.
Research paper analyzing fundamental problems with averaging LLM benchmark scores for model evaluation.
LLM-powered document extraction tool converting Bills of Lading and logistics documents to structured data.
Tool converting repository conventions into deterministic checks for AI agents to manage and enforce.
Celesto enables petabyte-scale persistent storage for AI agent sandboxes, useful for coding agents and large file handling.
Siplinx AI runs local LLM and speech-to-text models on-device for meeting transcription and note-taking without cloud.
Comparison guide of AI coding assistant subscription plans for 2026 with focus on developer workflows and value.
Agentic data engineering: using LLMs to generate SQL queries and automate analytics workflows. Practical workflow example.
Erlangchain: lightweight Erlang library for OpenAI/Anthropic LLM calls with tool-use and multimodal support, zero third-party dependencies.
MCP tool providing cross-runtime temporal knowledge graph memory for 7 AI coding agents, with constraint validation and context re-injection.
Essay on supervised vs unsupervised AI-generated code: examining human review in loop versus shipped-without-review code.
Security analysis of AI browser vulnerabilities: LLMs can be tricked via prompt injection to execute forbidden actions.
Liquid AI releases LFM2.5-230M, 230M-parameter model optimized for edge devices with fast inference for agentic tool-use workflows.
Ragit is a local RAG CLI tool for chatting with document folders using Ollama. Enables offline LLM document retrieval without API keys.
Rust-based unified agent substrate framework for governance and orchestration across multiple systems.
Open-source Aegize project implementing security layer for AI tools via identity, policy, and permissions controls.
Claude Code creator Boris Cherny discusses five job archetypes emerging in AI era: builder, operator, explorer, etc.
Label Imitation Game framework uses adversarial interrogation to prune hallucinations in foundation model pseudo-labels without standard thresholds.
arXiv paper introducing TheraJudge framework using multi-agent systems and human-aligned evaluation for mental health support with LLMs.
arXiv research on post-hoc detoxification of backdoored LLMs using curvature-guided module localization and low-rank repair.
Analysis of how human feedback shapes AI-generated Community Notes in X's crowd-sourced fact-checking system extended with collaborative AI features.
Budget-adaptive routing system for edge-cloud inference that skips weak model inference when offload budget allows, optimizing based on varying computational constraints.
Theoretical analysis explaining why on-policy distillation outperforms offline imitation learning with noisy expert feedback, with implications for language model training.
Loc2Repair modular evaluation framework for repository-grounded LLM repair systems isolating file-level issue localization as a key failure mode.
Analysis of deductive stereotyping failure mode in LLMs where models apply population-level statistics to individuals, with Fair-GCG mitigation approach.
Framework for using LLMs to drive decision-making behavior in virtual agents for emergency simulations, enabling believable autonomous agent behavior in interactive environments.
Knowledge distillation study from DeepSeek-R1 to Qwen2.5-7B using Chain-of-Thought training corpus from math competition problems with LoRA fine-tuning on Apple Silicon.
ADAPT method mitigating hallucinations in multimodal LLMs by aligning attention dynamics and using preference tuning to improve text-to-image cross-attention during generation.
Triospect framework for detecting AI-generated text by analyzing content, expression, and stylistic elements, robust against 17 attack types across multiple domains.
Training-free gated reranking method that uses model uncertainty to decide whether reranking few-shot examples improves LLM performance across NLU and machine translation tasks.
Modular vision-language-action robotics framework for autonomous agents performing complex indoor tasks from natural language instructions.
PruneGround method using spatial pruning to improve efficiency and accuracy in 3D visual grounding by focusing on relevant scene regions.
PPT-Eval benchmark with 120 PowerPoint tasks to evaluate computer-use agents on content creation and presentation editing scenarios.
Session-level RAG system reorganizing knowledge bases with co-occurrence clustering to cover multi-question user sessions beyond single-query retrieval.
Framework using LLMs to synthesize robotic actions from multimodal inputs including speech, gestures, and music for human-robot interaction.
ComplianceGate system routes LLM queries through multi-tier classifiers to enforce compliance and cost efficiency in regulated industries while protecting PII.
MIRTH framework for vision-language-action robotic agents addressing temporal understanding, reasoning gaps, and inference efficiency in physical control tasks.
Unified offline reinforcement learning framework for nonlinear multi-objective optimization capturing complex trade-offs like risk and fairness.
Dataset and evaluation of LLMs' ability to imagine moral alternatives beyond binary dilemmas in ethical reasoning tasks.
Evaluation framework using LLMs to detect stylistic copyright infringement in generated text beyond verbatim memorization detection.
One-step text-to-audio distillation framework using caption-only training data without paired audio for efficient generation.
Web-based toolkit for synthetic tabular data generation supporting Bayesian mixture models, diffusion, and latent-space generative modeling.
Inference-time self-improvement method for computer-use agents leveraging multimodal LLMs to autonomously improve from task failures.
Hierarchical memory bank method for online continual self-supervised learning from unlabeled data streams with memory constraints.
Method for post-training backdoor detection and trigger inversion in LLMs using class subspace orthogonalization.
AI-assisted workflow using artifact-transform language and LLM assistant to reduce visual analytics prototyping from months to hours.
3D trajectory guidance method for hierarchical vision-language-action models bridging planning and control in robot manipulation.
Study of probability calibration techniques to mitigate evaluator preference coupling bias in LLM agent feedback loops.
Visual reward-learning framework converting expert videos into dense rewards for training RL agents in long-horizon robotic manipulation tasks.
State-based fine-tuning approach for transformers using mixture-of-control for efficient parameter adaptation across blocks.