MP3: Multi-Period Pattern Pre-training for Spatio-Temporal Forecasting
MP3 pre-training method for spatio-temporal forecasting addressing temporal mirage problem in graph neural networks.
MP3 pre-training method for spatio-temporal forecasting addressing temporal mirage problem in graph neural networks.
Conformal Elo estimation framework for calibrating LLM-as-a-judge rankings while accounting for systematic biases and uncertainty.
Trajectory-based Quantization Sensitivity Score metric for post-training quantization using dynamical systems analysis.
Analysis of on-policy distillation in language and vision-language models examining sparsity and parameter geometry of updates.
Systematic review of AI and machine learning applications in library systems and information management.
Convergence rate analysis for partitioning classification under relaxed conditions with privacy considerations.
Theoretical analysis of generalized debiased Lasso estimator with stability principles and variable selection applications.
MirrorCheck framework detects adversarial attacks on vision-language models using text-to-image regeneration in multimodal settings.
Weakly supervised NLP pipeline for classifying diagnoses in Italian hospital discharge letters without manual document-level annotation.
Neuro-symbolic framework combining anomaly detection, symbolic reasoning, and RL for interpretable industrial digital twins.
UniversalRAG: Multi-modal retrieval-augmented generation over corpora combining text, images, videos. Unifies diverse modality RAG into single framework.
Reward-SQL: Text-to-SQL via RL with execution-aware reasoning and process-supervised rewards. Improves LLM complex query generation with database feedback.
Studies demographic bias in text-to-image models using LLM-based text conditioning. Shows implicit demographic assumptions even with unspecified attributes.
First LLM-based compiler for trapped-ion quantum computers. Fine-tunes LLMs to learn shuttling operations for quantum qubit layout-independent compilation.
Width pruning analysis of Llama-3.2 revealing trade-offs between knowledge retention and instruction-following. Shows GLU-MLP pruning dichotomy in model capabilities.
Method to detect lookahead bias in LLM economic forecasts via Lookahead Propensity statistic. Identifies information leakage in model training data.
DevRev: Addresses multi-tenant retrieval systems using query adaptation without full re-indexing. Leverages unlabeled query logs for domain adaptation at scale.
CuMA: Aligns LLMs with diverse cultural values using demographic-aware mixture of adapters. Addresses mean collapse and cultural sparsity in model alignment.
FusionRoute: LLM collaboration framework routing tokens to specialized models for improved performance across domains. Enables efficient multi-domain LLM inference.
Theoretical framework proposing symmetries as basis for defining interpretability in AI models. Argues existing interpretability definitions are untestable.
Benchmark for vision-language models combining fine-grained visual grounding with knowledge retrieval. Tests VLM capabilities on real-world high-resolution scenes.
Federated learning method for hierarchical systems using sign-based gradient compression. Machine learning research on distributed training optimization.
NeST: neuron selective tuning method for LLM safety that provides parameter-efficient and maintainable alternative to full fine-tuning.
Study of prospective memory failures in LLMs when formatting constraints conflict with concurrent task demands across 8,000 prompts.
Rabtriever: efficient rationale-based retrieval system using LLM-based generative rerankers with on-policy distillation for cross-encoding.
Analysis of feature computation budget's influence on per-instance algorithm selection effectiveness in black-box optimization.
AdaTKG: adaptive memory mechanism for temporal knowledge graph reasoning that maintains entity representations based on interaction history.
Unified framework for modeling structured flows combining source/sink behavior, cyclic dynamics, and topology-constrained transport.
Method for using LLM-guided priors in multi-objective Bayesian optimization with evidence-gating to calibrate LLM confidence to objective values.
Research on identifiable Markov Switching Models with instantaneous effects for modeling non-stationary temporal systems.
Patcher: defense method against backdoor attacks in LLMs that identifies and mitigates poisoned safety alignment without requiring attack details.
Empirical study on Direct Preference Optimization for fine-tuning LLMs, showing improved efficiency and competitive performance with simplified training pipeline.
Comparative study of autoregressive LSTMs, VAEs, and GANs for generating Bach-style symbolic piano music using MIDI data.
Discusses AI-native software engineering practices and team structures for the AI coding era.
Free tool that predicts whether an LLM can run on a user's GPU hardware.
Tool enabling Claude Code tasks to auto-resume after API quota resets, maintaining state in HANDOFF.md with automatic retry scheduling.
Memory layer for Claude Code and Codex CLI that stores coding conventions and project preferences to improve agent task execution.
OpenAI announces partner network to help enterprises identify use cases and integrate frontier models with existing systems at scale.
LLM agents are being used in security exploitation workflows to automatically find and exploit complex vulnerabilities in systems like Salesforce.
Anthropic's Constitutional AI uses RLHF with AI-generated feedback (RLAIF) to train Claude, raising questions about how AI welfare and alignment are measured.
graphCTX tool gives AI agents persistent memory of repo context (commands, conventions, decisions) so developers spend less time re-explaining and more time shipping.
HumanizeHub marketplace connects AI-generated markdown content with human editors for humanization and styling in protected workspace.
Essay comparing metaphors for AI-assisted coding: surgeon-assistant model versus commodity-on-meter, discussing delegation and responsibility.
Founder story of running unlimited $6/month LLM provider on consumer GPUs, discussing sustainability and AI agent reliability.
Agent Gate is a deterministic CI/CD firewall that validates AI-generated pull requests without executing untrusted code or LLM calls.
Bastion provides isolated Linux VMs for running multiple background coding agents in parallel without conflicts, fully self-hosted.
Open-source desktop agent that automates computer tasks, streams actions, and supports custom LLMs via BYOK.
ClawMoat provides runtime containment and security sandboxing for AI agents with tool use capabilities.
Community discussion comparing low-cost Chinese LLM models: DeepSeek, MiniMax, Qwen, GLM.
Lime 2.0 provides cryptographic JWT verification for AI agents without server calls, preventing DDoS economically.