UltraLIF: Fully Differentiable Spiking Neural Networks via Ultradiscretization and Max-Plus Algebra
UltraLIF: fully differentiable spiking neural networks using ultradiscretization and tropical geometry instead of surrogate gradients.
UltraLIF: fully differentiable spiking neural networks using ultradiscretization and tropical geometry instead of surrogate gradients.
SWE-MiniSandbox: lightweight container-free environment for scalable reinforcement learning training of software engineering agents.
Credal Concept Bottleneck Models: framework separating epistemic and aleatoric uncertainty in predictive models for reliable decisions.
Mathematical framework for linear representation hypothesis: analysis of how many neurons encode features in language model layers.
HiFloat4: block floating-point data format for efficient language model inference with 4.5 bits per value on average.
CryptoAnalystBench: benchmark for LLM agent failures when reasoning over complex multi-tool outputs and large document sets.
Predictive Associative Memory: neural architecture using temporal co-occurrence for memory retrieval instead of similarity-based approaches.
Security threat modeling analysis of AI agent communication protocols (MCP, A2A, Agora, ANP) identifying vulnerabilities in multi-agent interaction systems.
Co-design study exploring how theory-of-mind capabilities could be integrated into everyday user-facing AI products and services based on practitioner input.
Reformulates multi-objective combinatorial optimization as online learning over decomposed spaces using adaptive expert-guided sequential construction.
Analyzes whether self-referential vocabulary in LLM introspection reflects internal computation or confabulation using activation analysis and pull methodology.
Proposes bootstrapping-based regularization to reduce prediction instability in clinical deep learning models for patient care.
Improves LLM reasoning by identifying critical tokens and using paraphrastic probing with consistency verification to reduce hallucinations and error accumulation.
Proposes DiffuTruth framework using diffusion model likelihoods and non-equilibrium thermodynamics to detect LLM hallucinations without supervision.
Uses distillation to convert Transformers into SSM hybrids by preserving attention heads that perform in-context retrieval and distilling remaining layers.
Proposes efficient steering methods for unconditional diffusion models without gradient-based guidance during inference, enabling faster controllable generation.
Investigates whether a single learned representation can optimize performance across multiple reward functions in reinforcement learning using large model priors.
Examines how social media engagement signals create credibility proxies that may delay verification and lock in narratives about agentic AI systems before direct evaluation.
Training method unifying supervised fine-tuning with entropic objectives using gradient-based token weighting to resolve plasticity-stability tradeoffs in SFT.
LLM-based approach for detecting critical errors in machine translation including factual distortions and intent reversals for safety and reliability.
Denoising diffusion-based generative approach to learning-to-rank that models joint distribution over feature vectors for information retrieval.
Compiler-guided inference-time adaptation enabling GPT-5 to improve programming performance in low-resource languages like Idris through iterative feedback.
Theoretical framework (KPM) for studying persuasive interactions between generative social agents and humans based on knowledge and adaptive communication.
Roofline-based benchmarking framework for characterizing small language model performance on resource-constrained edge hardware and heterogeneous platforms.
Framework for fact-level attribution in multimodal LLMs enabling verification of multi-step reasoning and grounding outputs across heterogeneous input sources.
Differential privacy and communication efficiency approach for LLM split inference using stochastic quantization and soft prompts on resource-constrained devices.
Framework proposing six autonomy levels for GUI agents to standardize capability assessment and benchmark progress toward trustworthy software interaction.
Adaptive milestone reward mechanism for reinforcement learning-based GUI agents that resolves temporal credit assignment in long-horizon tasks.
Proactive defense method against attribute inference attacks in LLMs using word-level precision anonymization to prevent privacy breaches.
Krause Attention mechanism for transformers that prevents representation collapse and attention sink phenomena through bounded-confidence dynamics.
Training approach for reasoning models that learn from unverifiable data without relying on external verifiers or human-annotated reasoning examples.
Analytical Search framework for trend analysis and causal assessment using LLMs at corpus scale beyond standard RAG.
Analysis showing gradient compression hurts generalization in federated learning; proposes sharpness-aware minimization remedy.
ABot-N0 Vision-Language-Action foundation model unifying embodied navigation tasks via LLM reasoning and flow matching.
ArGEnT transformer architecture for learning solution operators across varying geometries in scientific machine learning.
ScalSelect training-free data selection method for efficient visual instruction tuning of vision-language models.
DMind-3 system for secure financial transaction execution in Web3 using edge-local-cloud AI stack with latency constraints.
LoRA-based parameter-efficient tuning for deploying LLMs on edge devices for malware detection with memory constraints.
SToRM: supervised token reduction method for multimodal LLMs in autonomous driving with human-vehicle interaction.
Theoretical framework for offline reinforcement learning in cyclic MDPs with heterogeneous stage-specific dynamics.
PatientHub: unified modular framework for simulating patients with LLMs for training counselors and therapeutic assessment.
DRACO: benchmark of 10 domain deep research tasks from real-world usage patterns for evaluating complex research capabilities.
ANML: machine learning framework that weights training samples by quality factors for improved robustness with specialized expert data.
TabSieve: select-then-predict framework that makes in-table evidence selection explicit for tabular prediction tasks.
LLM-driven 3D scene generation for agricultural simulation environments with domain-specific reasoning and verification mechanisms.
Strategy for adapting vision-language models to e-commerce data with multiple images, structured attributes, and noisy information.
AmbiBench: benchmark for mobile GUI agents that require clarification and interaction to understand ambiguous user instructions.
FLCOA framework analyzing cooperation breakdown in LLM-based multi-agent systems under communication delays and computational constraints.
MiniCPM-SALA: 9B hybrid LLM architecture combining sparse and linear attention for efficient long-context modeling.
Accelerated Prompt Stress Testing framework for evaluating LLM safety under repeated inference with identical or similar prompts.