FedAttr: Towards Privacy-preserving Client-Level Attribution in Federated LLM Fine-tuning
Privacy-preserving attribution methods for detecting data ownership in federated LLM fine-tuning using watermarking techniques.
Privacy-preserving attribution methods for detecting data ownership in federated LLM fine-tuning using watermarking techniques.
Unified self-distillation framework for LLMs without external teachers, addressing free-form trajectories and task-dependent correctness.
Benchmark study evaluating multimodal domain generalization methods to determine if performance gains reflect real progress or inconsistent evaluation protocols.
Retrieval-augmented agent framework for navigating large knowledge bases with intelligent query reformulation and evidence gathering.
Defense mechanism against backdoor attacks in federated learning using gradient analysis to detect and mitigate model poisoning.
Training method teaching metric distance relationships to discrete autoregressive LLMs for improved performance on numerical and spatial tasks.
Cross-domain Integrated Gradients method for generating saliency maps across multiple domains for time series model interpretability.
Theoretical analysis of depth's role in generalization using state-transition model separating implementation, approximation, and statistical error.
Position paper arguing constraints should replace fixed penalty methods in deep learning for trustworthy AI systems.
Provably safe reinforcement learning framework integrating analytic gradients with safety safeguards for autonomous robot deployment.
Survey on federated foundation models applied to recommendation systems while preserving privacy using decentralized learning.
Mechanistic interpretability method using attribution-guided pruning to discover and correct circuits in small-scale LLMs.
Variational formulation of Kolmogorov-Arnold Networks addressing hyperparameter selection for basis function count via Bayesian approach.
Reset Replay technique for sample-efficient LLM post-training via RL and preference optimization, addressing primacy bias and plasticity degradation.
Multi-objective instruction-aware RL for procedural content generation, improving controllability under complex natural language guidance.
Controlled study of multi-step reasoning in LLMs using cellular automata framework, showing models fail extended prediction despite learning local rules.
Synthetic data augmentation for conformal counterfactual inference, improving prediction intervals under treatment imbalance.
Formulates self-attention as robust state estimation with linear SDE model, matching standard attention complexity under isotropic noise assumptions.
Defense mechanism against membership inference attacks on diffusion models via higher-order Langevin dynamics.
Addresses extrapolation error in off-policy RL using friction analogy, representing replay buffer as action manifold with tangential/normal components.
Forensic analysis of synthetic data failure modes in model-based RL, diagnosing and solving issues in MBPO policy optimization.
Compressed KV-cache distillation enabling latent reasoning in LLMs, reducing computational costs while internalizing chain-of-thought.
Theoretical analysis of Reinforcement Learning with Verifiable Rewards (RLVR) training dynamics via gradient gap metric, explaining empirical success.
Examines cross-sample data memorization risks in federated learning of LLMs, adapting fine-grained detection methods from centralized settings.
Analysis of 129 LLM prompt datasets (>1.22TB) with linguistic patterns and downstream tasks: filtering, classification, quality prediction.
Gradient-free framework for test-time adaptation under distribution shift via latent subspace search, enabling efficient edge deployment.
Theoretical analysis showing learning causal parent sets is suboptimal for regret minimization in causal bandits with unknown structure.
Comprehensive review of Kolmogorov-Arnold Networks literature, clarifying relationships with MLP theory, kernel methods, and Kolmogorov superposition theorem.
Hybrid approach combining learned neural operators with iterative solvers to accelerate PDE computations while maintaining reliability outside training distribution.
Theoretical analysis proving high entropy regularization in Dec-POMDPs leads to convergent, symmetry-equivariant policies across different initializations.
Anthropic announces increased compute capacity and usage limits for Claude Code and Claude API through SpaceX partnership.
jj diff review tool integrated with AI agents for code review automation.
Sley is a new agent-native structural programming language designed for AI agents with deterministic edits and graph-aware tooling. Open source under Apache-2.0.
Commentary on AI startup failures, LLM wrappers becoming obsolete as foundation models offer direct functionality.
Security vulnerability in Claude Code sandbox allowing symlink escape that bypasses file access restrictions.
AnamDB: Rust neurosymbolic database engine integrating probabilistic neural perception with deterministic Datalog reasoning. AI-native differentiable architecture.
Carnegie Mellon, MIT, Oxford, UCLA study on cognitive impact of AI chatbot use. 10-minute exposure affects problem-solving.
Rust WebSocket relay for WordPress 7.0 Yjs CRDT collaborative editing, replacing HTTP polling with realtime events.
AniTroves anime database with LLM-based conversational discovery interface replacing traditional tag-based search.
Go tool providing local control plane for AI coding agents, managing package installs, shell execution, and API calls with policy enforcement.
Discussion of copyright and licensing issues for AI-assisted contributions to CPAN open-source modules, documenting major projects' policies.
Cost analysis of GPT-5.5 pricing and token economics compared to GPT-5.4 with 2x price increase and lower verbosity.
Conceptual argument for IDEs to evolve into programmable operating systems for AI, supporting agents, terminals, logs, tests, and long-running sessions beyond traditional file-based workflows.
Open source repository with 500k+ stars containing step-by-step guides for implementing technologies from scratch, maintained by CodeCrafters for learning programming fundamentals.
Selvedge: MCP server providing long-term memory for AI-coded codebases by logging reasoning behind code changes made by Claude Code, Cursor, and Copilot in SQLite.
Armorer: secure local control plane for AI agents managing lifecycle with Docker isolation, unified UI/CLI for monitoring, and dependency management for agent setup.
Work-in-progress implementation of AI-agentic surface called Surface, based on book 'The Evolution of Software Scale', with GitHub repositories for desktop and agent components.
OpenTelemetry exporter for incident alerting that integrates trace data with PagerDuty, OpsGenie, and Slack notifications.
Developer built ScreenKite product using AI, spending $130K+ on tokens; claims to match/exceed Screen Studio with faster 4K export.
Empirical evaluation showing AI trading systems underperform in contests, with inconsistent decision-making and excess trading.