Trust and Reliance on AI in Education: AI Literacy and Need for Cognition as Moderators
Studies how students' trust in AI assistants affects appropriate reliance and critical evaluation in educational settings.
Studies how students' trust in AI assistants affects appropriate reliance and critical evaluation in educational settings.
SelfGrader detects jailbreak attacks on LLMs using token-level logits without latency overhead or text generation randomness.
Evaluates formal reasoning capabilities of LLMs using Chomsky hierarchy benchmarks to systematically assess their understanding of computational complexity.
ChatSVA uses task-specific LLMs to generate SystemVerilog Assertions for hardware verification, addressing low functional accuracy and data scarcity.
Token-budget-aware pool routing optimizes vLLM inference by routing requests based on estimated token requirements, reducing resource waste and KV-cache failures in production deployments.
Grid2Matrix benchmark revealing vision-language models fail at exhaustive visual detail capture in controlled color-grid tasks.
A-IO framework for adaptive LLM inference on memory-bound NPU platforms addressing memory bottlenecks during decoding.
RoboLab: simulation benchmark for evaluating generalist robotics policies with minimal domain overlap between training and evaluation.
Research showing explicit sentence boundaries improve LLM capabilities by leveraging natural language structure.
Research on KL divergence stability under Gaussian perturbations beyond Gaussian families for OOD detection.
THEIA: modular neural architecture learning Kleene three-valued logic end-to-end without symbolic solvers.
Studies how LLMs simulate social behaviors in debate scenarios, analyzing agreement drift and network effects.
Identifies gradient entanglement problem in generalized category discovery and proposes energy-aware gradient coordinator.
MixAtlas method optimizes data mixture composition for multimodal LLM midtraining with benchmark-targeted recipes.
Formal methods combining XAI and neural network verification for trustworthy explanations with mathematical guarantees in safety-critical domains.
Graph-based fraud detection using GNNs with dual-path filtering to handle relation camouflage and heterophilic structures in fraud graphs.
TOPCELL framework using LLMs to optimize transistor topology in standard cell design, reformulating high-dimensional search as guided generation.
Policy learning under adversarial and exogenous environments with regret and safety violation guarantees for multi-agent decision-making systems.
Counterfactual routing method for sparse Mixture-of-Experts addressing hallucinations by activating dormant experts on long-tail knowledge.
Metric-Aware PCA unified framework for scale-invariant representation learning with spectral bias control via metric matrix parameterization.
Calibrate-Then-Delegate framework for LLM safety monitoring using model cascades with risk and budget guarantees via correctable error prediction.
GUI-Perturbed framework revealing systematic brittleness in GUI grounding models through domain randomization of visual scenes and instructions.
Behavior-regularized RL for LLM finetuning using value gradient flow to prevent value over-optimization from out-of-distribution extrapolation.
Contribution weighted group relative policy optimization for training LLM-based search agents with improved credit assignment and value estimation.
Tensor network methods from quantum physics integrated into ML models to mitigate exponential complexity via compressed representations.
Sample complexity bounds for best-arm identification in autonomous reasoning with LLM-based surrogate models and systematic evaluation bias.
Modular architecture for continual learning using task-specific experts and gatekeeper routing to prevent catastrophic forgetting.
VLM-to-IRL framework for esports player selection based on learned tactical decision-making patterns.
Theoretical analysis of multi-layer state-space models' expressive power and role of chain-of-thought in compositional tasks.
Class-incremental concept bottleneck model addressing catastrophic forgetting in continual learning with interpretability.
LLM embeddings applied to clinical records for early prediction of post-traumatic epilepsy from traumatic brain injury.
Framework leveraging LLM outputs as auxiliary data for parameter estimation in operations management tasks.
Autonomous agent framework using survival analysis for DeFi liquidation prevention and risk management.
ConfLayers: Adaptive layer-skipping self-speculative decoding speeds LLM inference via confidence-based pruning.
ELMoE-3D: MoE-based speculative decoding technique improving LLM serving efficiency on memory-bound hardware.
Zeroth-order optimization stability analysis for gradient-free learning and memory-efficient model fine-tuning.
MeanFlow: Few-step flow models replace diffusion for efficient policy representation in online RL.
Constraint-based pre-training paradigm enabling model scaling flexibility across varying deployment sizes.
Constrained policy optimization framework allocates test-time compute dynamically across LLM reasoning tasks for efficient inference.
PASS@(k,T) metric analyzes whether RL genuinely expands LLM agent capabilities beyond sampling, focusing on agentic tool use.
Proposes xFODE framework combining fuzzy and neural ODEs for interpretable nonlinear system identification with physical meaning.
Evaluates LLM juries against expert clinician panels for scoring medical diagnoses and clinical reasoning on real hospital cases.
Proposes Rejection-Gated Policy Optimization replacing importance sampling with learned acceptance gates for improved policy gradient updates.
Applies combinatorial satisficing bandits to mmWave beam and rate adaptation in multi-user wireless systems using ACK/NACK feedback.
Introduces LongAct framework leveraging intrinsic activation patterns in LLMs for long-context reinforcement learning and improved reasoning capabilities.
Proposes dynamic attention mechanism for sparse autoencoders to improve interpretability of foundation model activations by optimizing sparsity levels per neuron.
Uses LLM predictions as pseudo-observations to improve contextual bandit algorithms during cold-start, combining LinUCB with calibration-gated weight injection.
Knowledge distillation method (DLink) for compressing EEG foundation models while preserving task-relevant semantics for embedded BCI systems.
Research examining disagreement among fairness metrics in ML evaluation and reliability of demographic fairness assessment.
Research on verifiable gradient inversion attacks in federated learning with certification methods for successful privacy breaches.