Formally Verifying AI-Generated GPU Kernels
Gimlet uses formal verification to validate AI-generated GPU kernels. Research system complementing numeric tests with theorem proving.
Gimlet uses formal verification to validate AI-generated GPU kernels. Research system complementing numeric tests with theorem proving.
Analysis of overly-cautious code generation patterns in LLMs and their impact on code quality.
R3 is a local web UI for structured code review of AI agent outputs. Tracks feedback across document sections.
Research on proactive enterprise agents using Context Graphs for RAG systems that surface relevant information without explicit user queries.
Research on adversarial social epistemology for interactive systems with humans and LLMs, addressing information distortion and trust in testimonial chains.
Survey of LLMs for clinical reasoning and medical applications, connecting clinical practice with computational methods using Miller's Pyramid.
Framework for assessing alignment of LLMs in mental health applications, addressing safety risks beyond acute harms.
Infinity-Parser2 is a multimodal model using controllable data synthesis and reinforcement learning for document parsing with open-sourced training data.
VectorizationLLM: Custom LLM for teaching vectorization and mathematical analysis concepts in MATLAB for electrical engineering coursework.
Application of agentic AI, RAG, and multi-agent systems to actuarial underwriting, spanning rule-based automation to LLM-driven planning and tool use.
Feedback Manipulation Regularization: Offline agent alignment technique combining human demonstrations and feedback for improved imitation learning.
Nigeria Machinery Dataset: 89 industrial machinery records from Nigerian manufacturing/oil sectors with domain-grounded reasoning layer for low-resource analysis.
Persona Cartography: Method for decomposing and controlling LLM personalities using OCEAN framework and low-rank adapters on weight space.
Agentic Neural Architecture Search: Method bridging LLM-driven open-ended architecture design with NAS-driven optimization for automated model search.
Concretized Proposition Prompting (CPP): Framework improving LLM reasoning by explicitly grounding propositions, balancing compositionality and knowledge.
Harness engineering approach for productionizing LLM agents: moving behavior from prompts to deterministic code, schemas, and validation for auditability.
AegisDx: Safety-oriented framework for AI-assisted medical diagnosis using coordinated LLM components with structured reasoning and verification mechanisms.
Empirical analysis showing agreement among LLM judges does not reliably indicate correctness, challenging assumptions in ensemble LLM evaluation systems.
Study on adversarial persuasion attacks against chain-of-thought monitoring in LLM agents, demonstrating vulnerability of safety mechanisms.
CausalDS: Benchmark for evaluating causal reasoning in LLM-based data-science agents combining abstract reasoning with tool use on realistic data analysis tasks.
Neurosymbolic methodology integrating answer set programming with energy-based models for joint optimization, reasoning, and learning with background knowledge.
Overthinking technique: Using reasoning task vectors to amplify latent reasoning in language models for improved auditing and elicitation of hidden information.
ASMR: Multi-agent framework for automatic schema discovery from ship maintenance reports using field extraction and semantic analysis agents.
Mathematical formulation of slow thinking and active perception in cognitive systems, with applications to training reasoning-capable large language models.
ZendoWorld: A benchmark environment for testing AI agents on visual concept induction, hypothesis formation, and active experimentation through game observation.
Framework for multi-teacher knowledge distillation where frontier LLM teachers compete via execution-based judge, then collaborate to create verifiable curriculum for training coding student models.
MentalHospital benchmark for evaluating LLM performance on complete psychiatric clinical encounters beyond isolated tasks, covering interviewing, examination, assessment, and planning.
Knowledge distillation study measuring performance of sub-1B on-device models trained from 8B reasoning teacher for structured text extraction tasks, analyzing per-subtask capability transfer.
PolyUQuest framework for structure-aware RAG over web content using heterogeneous graphs that preserve HTML structure, DOM hierarchy, and entity relations for improved retrieval.
PredicateLongBench benchmark systematically evaluates LLM long-context capabilities across difficulty axes, addressing limitations of existing benchmarks like NIAH that only measure average-case performance.
Proposes psychological competence as missing evaluation dimension for AI systems used as advisors, coaches, and tutors.
LSTM framework for vehicle intention prediction in intersections with comprehensive ablation analysis for autonomous driving.
Benchmark exposing blind spots in multimodal models on tasks humans find trivial, evaluating failure modes beyond established benchmarks.
Semantic-aware discrete diffusion model generating realistic human mobility data while preserving privacy and discrete events.
One-shot federated learning approach using analytic visual prompt tuning to minimize communication bandwidth in edge deployment.
Studies mechanistic reasons why memorized knowledge fails to generalize in LLM fine-tuning, characterizing the knowing-using gap.
Multi-agent framework integrating game theory and Bayesian principles to reduce LLM hallucinations in rule-based scientific domains.
Benchmark evaluating vision-language models on nutrient reasoning and personalized health advice addressing information asymmetry in food systems.
Applies JEPA-style predictive learning to JA4 network fingerprints using Transformer model on cybersecurity data.
Drift-aware temporal graph framework capturing semantic evolution in biomedical text for improved retrieval and knowledge discovery.
AI-guided stimulus discovery and generation to optimize facial emotion perception studies in autism research.
XAI-guided adaptive fusion method combining unimodal and cross-modal experts for emotion and sentiment recognition.
Clinical-reasoning LLM for hepatocellular carcinoma risk stratification and treatment guidance from EMR narratives.
Analyzes real patient-chatbot conversations to understand communication patterns and develops patient simulator for evaluating health LLMs.
Multi-agent marketplace simulation studying formal mechanisms to prevent defection and maintain market stability among self-interested agents.
Physics-constrained benchmark for evaluating trustworthiness of autonomous agents in decentralized energy markets considering task performance and exploitation risks.
Addresses behavioral state decay in long-horizon AI agents through proactive memory mechanisms to surface decision-relevant information across expanding trajectories.
Analyzes quantization effects on LLMs beyond accuracy metrics, introducing correctness agreement to measure behavioral changes in quantized models.
Proposes symbolic workflow model for LLM-mediated applications using tool use, retrieval, branching, and checkpointing with Lisp-inspired conceptual framework.
Visual question answering benchmark for incident-centric dashcam understanding in autonomous driving using vision-language models.