Show HN: Restailor – open-source AI job fit/resume tailor/job tracker
Open-source resume and job application tool using LLM integrations from multiple providers for tailoring and job tracking.
Open-source resume and job application tool using LLM integrations from multiple providers for tailoring and job tracking.
Comparison of three methods (Ablation, Heretic, Obliteratus) for removing refusal behaviors from LLMs.
AI platform for converting clinical medical text to ICD-10 and SNOMED-CT codes using multi-layer ML architecture with custom models.
System automating data product improvement using specialized AI agents to generate supporting assets and example queries without manual expertise.
Structured memory system for GUI agents combining trajectory-based and learned memory to handle long-horizon workflows and interface diversity.
HEAL method for distilling reasoning from large reasoning models into smaller models using hindsight entropy to overcome teacher ceiling limitations.
TRACED framework evaluates LLM reasoning reliability through geometric kinematics measuring progress and stability rather than scalar probabilities.
Framework using imprecise probabilities to capture and verbalize higher-order uncertainty in LLMs beyond classical probabilistic methods.
Lightweight hybrid framework combining LLMs and graph attention networks for resource-constrained game-playing decision-making.
Training dataset addressing instruction hierarchy in frontier LLMs to defend against jailbreaks, prompt injections, and agentic attacks.
Approach using reward-free self-finetuning agents for adaptive RAN slicing control in network systems overcoming context window limitations.
Meta-evaluation framework using vision-language models to audit and evaluate autonomous computer-use agents at scale beyond static benchmarks.
Empirical study examining whether LLM alignment for moral reasoning requires diversity-seeking algorithms versus reward-maximizing policies using RLVR methods.
Framework for extracting actionable learnings from LLM agent execution trajectories to improve future performance and reduce repeated errors.
FAME proposes formal abstract minimal explanations for neural networks using abstract interpretation, scaling to large models with reduced explanation size.
DxEvolve is a self-evolving diagnostic agent mimicking clinician cognition through interactive deep clinical research with auditable improvement mechanisms.
Framework for building domain-expert AI agents through conversational knowledge crystallization, shifting from code-first/prompt-first approaches.
PharmGraph-Auditor system for safe medication verification using LLMs with knowledge graphs for traceability and complex reasoning in healthcare.
Research addressing calibration degeneration in reinforcement learning from verifiable rewards by decoupling reasoning and confidence in LLM training.
Parameter-efficient fine-tuning approach enabling single LLM to handle multiple code analysis tasks with reduced computational cost.
Research on LLM unlearning through reasoning to remove undesirable knowledge while mitigating safety, copyright, and privacy issues.
AraModernBERT adapts ModernBERT architecture to Arabic with transtokenized embedding initialization and native 8,192 token context support.
Research on efficient MoE model inference on edge devices using speculative activation utility and lookahead sensors for memory management.
Empirical study investigating whether LLMs exhibit Dunning-Kruger effect patterns, analyzing confidence calibration across state-of-the-art models.
Research paper quantifying hallucination rates in LLMs on medical textbook QA tasks, measuring factually incorrect claims against fixed evidence sources.
Demonstration optimization method using evolutionary algorithms for chain-of-thought feature transformation in data-centric AI tasks.
Pipeline for mechanistic interpretability combining activation patching with natural language explanations of LLM internal circuits.
System Hallucination Scale (SHS): human-centered measurement instrument for evaluating factual unreliability and hallucination behavior in LLMs.
Two-stage NDA analysis system using LLaMA for clause extraction and transformers for contract clause classification.
Framework for building domain-adapted LLM conversational systems using fine-tuning, RAG, and evaluation methodologies for institutional deployment.
CEI benchmark with 300 validated scenarios for evaluating pragmatic reasoning and contextual inference in LLMs.
Analysis of adjective-noun compositionality in LLMs using prompt-based and representational methods revealing performance-internal state divergence.
Study comparing human-in-the-loop approaches with chain-of-thought prompting for behavioral interview evaluation using LLMs.
Evaluation of offline LLM capabilities for Turkish heritage language education focusing on pedagogical safety and robustness.
Clinical evaluation of GPT model generations for psychological safety across emotionally challenging conversational scenarios.
Automated evaluation framework for assessing LLM machine translation quality from Mandarin Chinese to English.
Retrieval-augmented assistant for unmanned aircraft safety assessment and regulatory compliance automation.
Dataset creation method using Wikidata to detect sociocultural bias in LLMs for Latin American languages and contexts.
Benchmark platform for evaluating LLM performance on end-to-end spreadsheet generation with explicit and implicit user constraints.
Arabic medical text classification system using fine-tuned AraBERTv2 encoder benchmarked against multilingual and Arabic-specific models.
Method for personalizing LLM alignment to diverse user preferences beyond single global objectives using group-relative policy optimization.
Automated red teaming framework for generating multi-modal adversarial conversations to test LLM robustness with expansion techniques.
Research on measuring and reducing safety refusals in military LLMs to enable accurate information provision in combat scenarios.
Study assessing cognitive biases in LLMs (virtuous victim effect, halo effect) relevant to judicial decision support applications.
DeliberationBench benchmark for measuring persuasive influence of LLMs on user beliefs using deliberative polling methodology.
Policy analysis defining 'AI model' and 'AI system' through systematic review of 896 papers and 80+ regulatory/technical documents.
RedFuser automatic operator fusion framework for cascaded reduction operations in attention mechanisms on AI accelerators.
Policy paper on legal and accountability frameworks for identifying and governing autonomous AI agents in economic systems.
Linux kernel module dmaplane for buffer orchestration and lifecycle management in high-performance AI data paths.
Benchmark study of LLM inference optimization on AMD Instinct GPUs for 235B-1T parameter models with architecture-specific deployment guidance.