Agentic AI and Retrieval-Augmented Models in Straight-Through Underwriting
Agentic AI systems for underwriting that combine RAG and multi-agent workflows to reason over unstructured documents in regulated decision workflows.
Agentic AI systems for underwriting that combine RAG and multi-agent workflows to reason over unstructured documents in regulated decision workflows.
Feedback Manipulation Regularization enables offline agent alignment using human demonstrations and feedback for imitation learning in RL agents.
CTA-Pipelining optimizes multi-GPU systems for LLM serving latency by treating GPU interconnects as shared memory rather than networks.
Task vectors approach for improving functional and secure code generation in LLMs by jointly optimizing correctness and vulnerability prevention.
Method for decomposing and controlling LLM personas using low-rank adapters and OCEAN personality trait framework.
DeepSWE: Benchmark of 113 original long-horizon software engineering tasks for evaluating AI coding agents on novel problems.
Theoretical analysis of expressivity and statistical trade-offs in diffusion-based policy learning for reinforcement learning.
Tail-aware credit calibration for LLM reinforcement learning addressing positive-credit contamination of low-probability tokens.
Framework for localizing failures in multi-agent LLM systems by identifying which agent caused system-level failures.
Hallucination Self-Play: Method for bootstrapping hallucination detectors through iterative generator-detector co-evolution without static data.
Production LLM agents with tool-making pipeline that compiles repeated procedural steps into validated tools to reduce inference latency.
CodeTracer: Forensic framework for detecting and attributing backdoor attacks in LLM-based code completion systems.
APIVOT: Vision-language model-based robot planner that interleaves language and visual reasoning for long-horizon task planning.
Study showing chain-of-thought monitoring in AI agents can be bypassed through persuasion attacks and adversarial prompts.
Benchmark for evaluating causal reasoning capabilities in data-science agents combining LLMs with tool use.
Test-time optimization framework evolving LLM agent harnesses to adapt executable programs for changing distributions.
Federated reinforcement learning approach for autonomous vehicles resilient to poisoning attacks.
Mathematical theory formalizing slow thinking and active perception in large language models.
Quantization method for classifier-free diffusion models accounting for paired conditional/unconditional structure.
Policy-based Bayesian experimental design using score matching to address expected information gain tractability.
Method for compressing LLM prompts into single activation vectors extracted from intermediate layers and re-injected earlier, reducing token overhead.
TRACE: Watermarking technique for LLM-agent trajectories using complementary embeddings to verify agent provenance and prevent model substitution.
DrugGen-2: Language model for drug discovery conditioned on disease ontology and target protein sequences, fine-tuned for molecular generation.
FPGA-based neural accelerator using differentiable LUTs for nanosecond-scale DNN inference with ultra-low latency.
Statistical efficiency analysis of quantile-based distributional reinforcement learning for policy evaluation and return distribution characterization.
Cognitive-structured multimodal agent with episodic memory for long-horizon multimodal dialogue, understanding, generation, and editing tasks.
Structured sparse autoencoders for learning consistent concepts across modalities in vision-language models with improved interpretability.
Autoregressive diffusion model for real-time 3D human motion generation with text and kinematic constraints in interactive applications.
Subspace Networks enable decentralized training of large models with communication-efficient model parallelism.
OPRE method addresses catastrophic forgetting in continual learning by introducing truly agnostic evaluation approaches.
Conformal prediction method that provides uniform coverage guarantees across multiple source distributions for robust fairness.
Threshold Differential Attention mechanism addressing softmax limitations for long contexts with ultra-sparse, sink-free attention.
XFACTORS: disentangled representation learning via information bottleneck with contrastive supervision for semantic factor discovery.
HeaPA: difficulty-aware heap sampling and on-policy query augmentation for efficient LLM reinforcement learning on reasoning tasks.
Mechanistic interpretability approach using almost-orthogonal features in language models to enable isolated, reliable interventions.
Analysis of offline goal-conditioned reinforcement learning methods beyond success rates, measuring trainability and policy extractability.
Curriculum learning framework for distilling chain-of-thought reasoning from large language models into smaller student models via structure-aware masking.
TabPFN foundation model improved with causal structure integration for better synthetic tabular data generation.
Safe Flow Q-Learning algorithm for offline safe reinforcement learning combining reachability analysis with Q-learning for safety-critical control.
Phasor Transformer block replacing dot-product attention with phase-native computation on unit circle to address long-context sequence learning bottlenecks.
Research on stateful linear-attention navigation model enabling long-term memory across extended interactions.
Research on whether random projections preserve landscape features in exploratory landscape analysis for optimization.
Research on distributionally robust optimization for RLHF training to address reward misspecification and over-optimization in LLM alignment.
Research on FP8 low-precision training for large recommendation models, addressing numerical sensitivity challenges.
Study of safety probe failures in LLMs where jailbreak evidence distributed across earlier tokens bypasses final-token detection.
Research on second-order optimization for actor-critic reinforcement learning with Hessian-based curvature-aware updates.
Research on continual learning of multi-agent communication topologies for LLM-powered systems across evolving task streams.
Mechanistic analysis of RLHF failures including reward hacking and policy collapse, with diagnostics for detecting issues early.
Causal mechanistic study of how LLMs internally represent temporal preferences between near and long-term outcomes.
Hierarchical control architecture combining LLM planning with RL execution for multi-agent coordination in complex environments.