CliffSearch: Structured Agentic Co-Evolution over Theory and Code for Scientific Algorithm Discovery
CliffSearch agent framework for scientific algorithm discovery combining LLM-guided search with structured evolution of theory and code.
CliffSearch agent framework for scientific algorithm discovery combining LLM-guided search with structured evolution of theory and code.
Mathematical framework analyzing what determines forecast skill in AI weather prediction, emphasizing training methodology over architecture.
PhoneticXEUS model for robust multilingual phone recognition trained on large-scale data with pretrained representations.
LLM-based recruitment tool identifying requisition-specific competencies through dynamic few-shot prompting and reflection.
Text-based harmonization approach using LLMs to unify multi-institutional EHR data without explicit schema standardization.
LLM-based approach to identify enterprise architecture debt indicators from unstructured documentation in organizations.
Framework combining vision language models with RL for dense reward generation in long-horizon robotic tasks to reduce manual reward engineering.
GenoBERT uses transformers for reference-free genotype imputation without ancestry bias.
HIVE framework for hierarchical pre-training of vision encoders integrated with large language models for vision-language alignment.
Transformer-based models for detecting software vulnerabilities in C/C++ using program slices.
MambaVoiceCloning uses state-space models and diffusion for efficient text-to-speech synthesis without attention layers.
Studies grokking in feature learning kernels via Recursive Feature Machine, showing data symmetry breaking is necessary for generalization.
Reverse-engineers gpt-oss-20b tool definitions from in-distribution calls and builds native harmony agent harness with open-source implementation.
Proposes decision-centric framework separating control decisions (answer, retrieve, tool use) from LLM generation in agent architectures.
EgoNav: Humanoid robot navigation system trained on 5 hours of human walking data using diffusion models and frozen DINOv3 backbone.
Shapley-guided approach using derivative-free optimization to repair DNNs affected by backdoors, adversarial attacks, and unfairness.
Studies policy gradient methods for multi-agent reinforcement learning in partially observable Markov potential games.
CheXOne: Vision-language foundation model for chest X-ray interpretation with explicit reasoning about visual evidence.
Introduces Uni-SafeBench, a safety benchmark for unified multimodal large models testing both understanding and generation capabilities.
Studies trade-off between pretraining corpus size and retrieval-augmented generation for language models under fixed data budgets.
CircuitProbe predicts reasoning circuits in Transformers from activation statistics in under 5 minutes, achieving 3-4 orders of magnitude speedup over brute-force methods.
Benchmarks State-Space Models (Mamba) against Transformers and BiLSTM for historical newspaper OCR, addressing quadratic complexity limitations.
Stochastic Attention inspired by connectome topology provides linear-time expressive attention mechanism.
Study shows multimodal LLMs fail at detecting 3D spatial inconsistencies across multiple views.
PARE framework simulates realistic user interactions for evaluating proactive AI agents and assistants.
StanceMoE uses mixture-of-experts for actor-level stance detection in geopolitical texts.
Dataset and analysis of autonomous coding agent contributions to real-world GitHub projects over time.
MyPhoneBench evaluates privacy compliance of mobile phone-use agents completing benign tasks.
Query-conditioned evidential keyframe sampling for efficient multimodal LLM-based long-form video understanding.
MoA-DepthCLIP adapts CLIP vision-language model for monocular depth estimation with parameter-efficient adapters.
PaperRecon framework evaluates quality and hallucination risks in papers generated by AI coding agents.
NARCBench for detecting multi-agent collusion using multi-agent interpretability on LLM agent activations.
S0 tuning zero-overhead adaptation of hybrid recurrent-attention models outperforming LoRA on code generation.
RELISH lightweight architecture for text regression with LLMs using iterative latent state refinement.
Survey on Graph Neural Network acceleration techniques across algorithms, systems, and customized hardware.
RobustRAG defense framework with certifiable robustness against retrieval corruption attacks on RAG systems.
Gradient-based hyperparameter learning via evidence lower bound objective from Bayesian variational methods.
Neural framework for learning conditional optimal transport maps using hypernetworks to generate adaptive transport parameters.
JUSSA framework uses steering vectors to improve LLM-as-a-judge reliability, detecting and mitigating sycophancy through honesty-promoting alternatives.
Binned semiparametric Bayesian networks for efficient kernel density estimation using data binning to reduce computational cost.
Double-Diffusion integrates ODE-prior with denoising diffusion models for spatio-temporal graph forecasting, balancing deterministic and stochastic components.
Klear-Reasoner model with long reasoning capabilities using gradient-preserving clipping policy optimization, with detailed training disclosures.
Knowledge component discovery in programming using representation learning on student code for personalized instruction systems.
Thompson sampling analysis for Sharpe ratio optimization in multi-armed bandit setting, addressing fractional objective with dependent mean-variance.
LSTM-based machine learning calibrator for agent-based epidemic models, learning inverse mapping from time series to SIR parameters.
EEG classification study comparing neural network architectures and optimizers across brain hemisphere frequency bands using TensorFlow/PyTorch.
Comprehensive survey of intrinsic dimension estimators under manifold hypothesis, reviewing theoretical foundations and comparing eight methods.
Analysis of weight constraints in linear smoothers for causal inference, balancing feature imbalance against parametric modeling assumptions.
Polychromic objectives approach to prevent mode collapse in reinforcement learning fine-tuning, preserving policy diversity during exploration.
Convergence analysis for decentralized SGD with high-probability guarantees, removing restrictive assumptions on gradient bounds and noise.