Beyond the Wrapper: Identifying Artifact Reliance in Static Malware Classifiers using TRUSTEE
TRUSTEE method to identify artifact reliance in ML-based static malware classifiers, revealing which features are semantic vs. spurious.
TRUSTEE method to identify artifact reliance in ML-based static malware classifiers, revealing which features are semantic vs. spurious.
POMDP framework for LLM agents to iteratively search and gather relevant context from large environments exceeding context window limits.
Interpretable and scalable LLM evaluation framework using Item Response Theory to model latent abilities and item characteristics beyond average accuracy.
Bayesian framework for physics-informed neural networks using functional priors for PDE-constrained inverse problems.
Self-Consolidating language models that write current context into model weights while incorporating new information and limiting catastrophic forgetting.
Visual feature-based world models using residual latent actions to predict future visual features instead of pixels for more efficient predictions.
Framework separating channel importance in vision networks into task relevance and local replaceability, proposing two-axis evaluation method.
Develops TRACE method for constructing valid conformal prediction regions using diffusion and flow matching models.
Proposes classification fields for infinite-depth recursive hierarchical clustering with fine-scale refinements from few examples.
Introduces AdaTKG with adaptive entity-level memory for temporal knowledge graph reasoning over evolving events.
Proposes RRCM framework for LLM-based recommender systems using ranking-driven retrieval over collaborative and meta memories.
Constructs Adversarial Empathy Benchmark to probe robustness of RL-trained empathetic language agents against adversarial user interactions.
Proposes Distillation through Reasoning Path Compression to improve consistency of teacher rationales when distilling reasoning into smaller LLMs.
Introduces MathlibPR benchmark for evaluating LLM-assisted pull request merge-readiness in Lean formal mathematics library.
Audits text embeddings using citation graphs of 3.58M papers, revealing disconnect between cosine similarity and conceptual relatedness in RAG.
Benchmarks attention transfer across 20 Vision Transformer teachers, finding it not universally effective across ViT families.
Develops closed-form dataset distillation method for linear probing on frozen pre-trained vision model encoders.
Proposes Three-in-One world model combining Deep Boltzmann Machine with task-specific heads for marketing intervention modeling.
System for toxicity detection in gaming chat using fine-tuned LLMs with LoRA and synthetic data augmentation across six toxicity classes.
Introduces proxy-analyzer framework using open-weight models to detect hallucinations in LLM outputs via internal activations.
Proposes Planning-after-Trial adaptive policy for test-time code generation that efficiently allocates compute based on problem difficulty.
Presents MIPIAD defense framework against multilingual indirect prompt injection attacks in RAG and tool-using LLM systems using ensemble learning.
Introduces structured role-aware policy optimization for multimodal reasoning in large vision-language models using reinforcement learning from verifiable rewards.
Proposes sparse random-feature neural networks with Krylov-based SVD for solving singularly perturbed ODEs with improved scalability.
Analyzes generalization bounds for trained Transformers using spectrum-adaptive methods to improve upon existing norm-based complexity bounds.
DoLQ: method for discovering ordinary differential equations combining symbolic regression with LLM-based qualitative evaluation.
Sparse autoencoder architectures (Crosscoders, Diff-SAE) for detecting backdoor attacks in language models via mechanistic interpretability.
Analysis showing mean-pooled cosine similarity is length-dependent; proposes length-invariant alternative for comparing neural representations.
Memory-efficient alternatives to error correction codes for protecting deep learning models against hardware faults.
Sparse autoencoders as lightweight detection mechanism for adversarial attacks on vision-language models.
Data selection framework for large multimodal models using incremental optimization utility ranking instead of LLM-as-Judge.
Reinforcement learning approach for distilling compact GUI agents that work on-device across platforms.
Conformal prediction framework for uncertainty quantification in object detection with finite-sample coverage guarantees.
Theoretical analysis of sample complexity in contrastive representation learning with dependent tuple sampling.
Framework for evaluating safety vs. capability in phone-use agents, addressing ambiguity in existing benchmarks.
Study showing post-training reduces LLM alignment with human behavior; introduces Psych-201 dataset for measuring behavioral alignment at scale.
Decentralized multi-agent pathfinding solver using learned local communication for scalable trajectory planning in robotics and logistics.
MAVEN multi-agent verification network with epistemic auditing for reasoning tasks, enabling intermediate verification in CoT traces.
Prefix consistency method for improving Chain-of-Thought reliability by using answer reproduction as verification signal.
Data valuation method using quotient semivalues to resist false-name manipulation in ML data attribution.
Game-theoretic analysis of differential privacy auditing when audited developers can strategically respond to audit queries.
FactoryBench benchmark evaluates time-series models and LLMs on industrial robotic telemetry along four causal levels.
Memory-efficient recurrent LLM architecture decoupling compute from memory in looped reasoning models with sublinear KV cache.
Research on LLM self-assessment using cognitive appraisal theory as alternative to confidence scores for performance prediction.
CADTestBench introduces first test-based evaluation benchmark for Text-to-CAD task using automated testing methodology.
arXiv: MatryoshkaLoRA - adaptive rank LoRA for efficient LLM fine-tuning without grid search. Parameter-efficient training.
arXiv: Quantum-inspired optimization for non-convex ML problems in high dimensions. Theoretical optimization approach.
arXiv research: Tool-calling in LLMs is linearly readable/steerable via internal activations. Tested 12 models, 77-100% steering accuracy.
First-order optimization methods for bilevel problems with minimax structures in upper and lower levels.
Global training method for Spiking Neural Networks via parameter reconstruction addressing surrogate gradient approximation errors.