Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
Research identifying capability ceiling in LLM RL post-training by reintroducing Markov state structure for improved alignment.
Research identifying capability ceiling in LLM RL post-training by reintroducing Markov state structure for improved alignment.
SuperKMeans, a k-means variant for clustering vector embeddings 7x faster than FAISS/Scikit-Learn on CPUs.
AgenticRS-EnsNAS addresses NAS validation bottleneck in ensemble-based production systems with decoupled architecture search.
Open-source framework for automated detection, segmentation, and severity estimation of lesions in coronary angiography images.
Continual learning method treating representation evolution as shared-manifold continuation to prevent catastrophic forgetting.
Hyperdimensional computing for federated learning on resource-constrained IoT edge devices.
RL algorithms for fine-tuning financial forecasters with supervised learning models, showing improved performance via transfer learning.
Theoretical framework interpreting diffusion model generation as out-of-equilibrium phase transitions with architectural analysis.
Research on spectral alignment in forward-backward representations for successor representation learning in continuous spaces.
λ-RLM: framework using λ-calculus to solve long-context problems in LLMs via recursive decomposition with verifiable control code execution.
Analysis of trojan horse attacks in deep forecasting models using insights from ESA competition, focusing on backdoor triggers in safety-critical applications.
Var-JEPA: variational formulation bridging JEPA and probabilistic generative modeling, showing their structural equivalence beyond rhetorical separation.
Method for conditioning protein sequence generation via stochastic attention using Hopfield pattern multiplicity to direct generation toward functional subsets.
Virtual study group of AI agents for knowledge discovery in gene ontology using LLMs and agentic AI to extract aging-related biological knowledge.
D-MMD: discrete moment matching distillation method for distilling discrete diffusion models to reduce sampling steps while maintaining quality and diversity.
KaCGM: causal generative model for mixed-type tabular data using Kolmogorov-Arnold Networks for interpretable structural equations and causal inference.
MAPLE: method for generating differentially private synthetic data from LLMs accessed via APIs, enabling DP fine-tuning without direct model access.
Significance-Gain BPE: statistical alternative to frequency-based subword tokenization for LLMs that measures pair cohesion rather than raw frequency for improved tokenization.
Novel LLM distillation framework using reinforcement distillation and explanatory inversion to transfer reasoning capabilities to smaller student models with better generalization.
Autonoma: hierarchical multi-agent framework for translating natural language instructions into robust multi-step workflows, addressing scalability and error propagation in workflow automation.
Modular summarization framework decomposing review processing into interpretable components for opinion extraction and clustering.
LLM-based approach combining daily financial news with stock embeddings for multi-stock price prediction.
LLM-based multi-agent framework for legal judgment prediction with interpretable reasoning and evolving case analysis.
Hypergraph neural network method for solving polynomial-objective integer programming problems with nonlinear constraints.
Study of agentic frameworks with LLMs for Verilog hardware code generation, comparing agent-based versus standard approaches.
Using LLM agents to automate discovery of membership inference attacks against machine learning models.
Dynamic routing system selecting optimal LLM from large model pools using fine-grained latent task discovery.
Principled framework for unsupervised domain adaptation in kernel GLMs under covariate shift using pseudo-labeling.
Analysis showing safety-focused defense training degrades LLM agent tool-use capabilities, revealing capability-alignment paradox.
Study on how vocabulary influences word-order learnability in transformer language models across languages.
Protein language model with reinforcement learning guidance for designing diverse AAV capsids for gene therapy.
Multi-modal language model agent trained with process-reward RL to generate vector sketches part-by-part.
Benchmark for evaluating vision-language models on medical photograph understanding tasks.
Reinforcement learning framework ensuring LLMs ground responses in evidence to reduce hallucinations.
Method for verifiable error bounds in physics-informed neural networks solving Lyapunov and HJB equations.
Test-time adaptive steering method improving object count accuracy in text-to-image diffusion models without retraining.
Framework improving long-horizon LLM agent planning using subgoals for complex web navigation and digital environment control tasks.
Theoretical analysis of scaling limits in generative models through Godel-Tarski-Lob framework examining capability growth bounds.
Neural network architecture that dynamically grows and prunes parameters during training for efficient image classification.
Analysis of CLIP vision-language model projectors to improve intra-modal alignment for image retrieval tasks.
Survey of deep learning architectures and objectives for modeling autocorrelation in time-series forecasting tasks.
Vision-language model for structured pathology report prediction treating multi-granular diagnostic output as primary objective.
Evaluation of test-time adaptation methods for facial expression recognition under natural cross-dataset distribution shifts.
Quantum architecture search method progressively growing parametrized quantum circuits for 3D point cloud classification tasks.
Theoretical analysis of adversarial learning with graph-factorized distributions using interpolative divergences and variational methods.
Self-supervised framework learning predictive wireless channel representations using JEPA world modeling on CSI trajectories.
Improves HAL semantic representations using attention-based pooling for text classification. Addresses information loss in mean pooling aggregation.
Semantic token clustering method for uncertainty quantification in LLMs without repeated sampling. Addresses computational overhead in reliability assessment.
Shows chain-of-thought faithfulness measurement varies by classifier, undermining single aggregate metrics. Evaluates 12 open-weight LLM models with 10K+ traces.
Demonstrates Claude Code autonomously executing high energy physics analysis pipeline including event selection, background estimation, and statistical inference.