EUISUH JEONG
AI engineer · computer vision · Staff Sergeant, ROKAF · CMU CS '22
Engineer by training, AI researcher by choice.
I'm an AI engineer and software engineer currently serving as a Staff Sergeant with the Republic of Korea Air Force, where I build computer-vision systems for runway integrity.
Before the service, I helped found aiXamine at QCRI — a platform that stress-tests language models against safety benchmarks. I'm a Carnegie Mellon CS '22 grad, with a minor in Mathematical Sciences.
I've been moving since I was three. Seoul, then a small town in the US, then back to Seoul, then India for secondary, then CMU, then Qatar for work, then home again. Cultures stack, like middleware. The interesting work happens in the seams.
Six years, three time zones.
AI Engineer · Staff Sergeant · Squad Leader
Led the squad that built and deployed an AI-driven runway pavement evaluation system at an active ROKAF airbase. Constructed a 231,347-image dataset — 52,800 real captures augmented with 178,547 alpha-blended synthetic images across 9 defect classes (SSIM 0.98163, FID 4.2145) — published as the ROKAF Runway Crack Dataset in KOSAP (Vol. 1, No. 2, Dec 2025) as first author. Co-designed the PCI scoring pipeline around YOLOv11, achieving 86.8% detection accuracy and a 98.2% reduction in manual assessment time; published in KOSAP (Vol. 1, No. 1, Aug 2025). Full project lifecycle as technical lead.
Research Engineer
Co-developed aiXamine — a black-box LLM safety evaluation platform with 40+ benchmarks across 8 security dimensions. Built the modular reporting + visualization architecture; evaluated 50+ models across 2K+ exams, surfacing vulnerabilities in GPT-4o, Grok-3, and Gemini 2.0. Also investigated backdoor Trojan attacks on code-focused LLMs (finetuning + susceptibility testing).
Software Engineer
Built a multi-channel notification system (SMS, email, push) for the consumer fintech app. Migrated payment processing to a compliant platform under regulatory scrutiny. Designed and shipped a Clubhouse-style waiting list + lottery system tied to FIFA World Cup Qatar 2022.
Teaching Assistant · 11-785 Deep Learning (PhD-level)
Planned and delivered lectures, recitations, and assignments to 350+ students in CMU's flagship deep-learning course. Mentored research projects and guided exploration of novel directions. Sample recitation on YouTube →
B.S. Computer Science · Minor, Mathematical Sciences · University Honors
Coursework concentrated in systems, machine learning, and applied math.
Things I built that went live.
Runway Evaluation System
Live · ROKAFDetects cracks and surface defects on airbase runways and computes PCI scores from high-res imagery. In operational use — 86.8% detection accuracy, two KOSAP papers published.
aiXamine
Live · publicA black-box LLM safety platform that runs repeatable exams across bias, robustness, jailbreaks, and other risk dimensions. I helped build the benchmark harness and reporting system.
Code-LLM Backdoor Poisoning
Research · completeInvestigated stealthy, trigger-based data-poisoning backdoors that coerce code-generating LLMs into emitting vulnerable source only when triggered — and where current vulnerability detectors fall short against them. Built on CVEFixesUtil, a tool I wrote for parsing the CVEfixes dataset across six languages.
CareRing
Competition entry · 2026AI wellness-call platform for elderly parents living alone: scheduled calls in natural Korean, health/emotional signals extracted from voice (response latency, speech rate, mood trend), a plain-language weekly report to the guardian, and keyword-based emergency escalation mid-call. Entry in the 2026 Air Force Startup Competition — I led the team and designed the AI call engine and voice health pipeline, building on prior emotion-analysis work at QCRI.
CrunchCut
Personal toolCross-platform video compression and audit tool. Drop a folder, get a live-streamed before/after size prediction per file with a recommendation level, then compress with a real-time progress view — no more re-running ffmpeg five times to find a setting that doesn't wreck quality. Ships as a native Tauri (Rust + React) app for Mac and Windows, bundling the original Python CLI as its backend.
Papers and conference work.
SurfaceGuardBench: How Far Do Lightweight Pattern Guards Get for Coding-Agent Tool Safety?
SurfaceGuardBench measures unsafe tool calls in coding-agent traces and tests lightweight, task-scoped guards before execution. It covers secret access, destructive shell use, network exfiltration, and instructions embedded in untrusted tool output.
RepOrbit: An Interpretable Event-Sourced Index for Repository Management
RepOrbit turns Git and GitHub history into reproducible repository timelapses and explainable Repository Management Quality Index (RMI) trajectories, connecting code changes with issues, reviews, CI, releases, and maintenance signals.
Music Tagging Graph Neural Network with Tag Labels
MTGNN is a graph neural network framework for music auto-tagging. It adapts ATGNN's graph-based audio-tagging idea to music by redesigning node generation around semantic and timbre features, then uses a CLAP-initialized Graph Transformer to model dependencies between tag labels.
ROKAF Runway Crack Dataset: Construction and Application of a Large-Scale AI-Based Runway Defect Detection Dataset
A large-scale runway defect dataset built from real airfield captures and synthetic augmentation for AI-based crack and surface-defect detection.
Deep Learning for Pavement Management System: Proposing an Automated Pipeline for Pavement Condition Index (PCI) Assessment
An automated pavement-management pipeline that combines runway defect detection with PCI scoring to reduce manual inspection work.
aiXamine: A Comprehensive Safety Evaluation Platform for Large Language Models
A safety-evaluation platform for large language models, covering bias, robustness, jailbreak, and other benchmark-driven risk checks.
Pictures.
Get in touch.
Open to research collaborators and post-service roles. Fastest reply by email.