Jae-Hong Lee

25 papers A* 4A 5B 2C 1Misc 3Journal 8Unranked 2
YearRankTypeTitle / Venue / Authors
2026 J jnl
IEEE Signal Process. Lett.
Ji-Hwan Mo, Dong-Hyun Kim, Jae-Hong Lee, Joon-Hyuk Chang
2026 J jnl
Comput. Speech Lang.
Jin-Seong Choi, Jae-Hong Lee, Joon-Hyuk Chang
2026 A* conf
AAAI
Chae-Won Lee, Jae-Hong Lee, Ji-Hun Kang, Joon-Hyuk Chang
2025 A* conf
ICML
Jae-Hong Lee
2024 J jnl
Robotics Auton. Syst.
Soo Ho Woo, Soon-Geul Lee, Jaehwan Choi, Junki Hong, Jae-Hong Lee, Haotian Xie
2024 A conf
INTERSPEECH
Mun-Hak Lee, Jae-Hong Lee, Do-Hee Kim, Ye-Eun Ko, Joon-Hyuk Chang
2024 A* conf
ICLR
Jae-Hong Lee, Joon-Hyuk Chang
2024 J jnl
IEEE Signal Process. Lett.
Chae-Won Lee, Jae-Hong Lee, Joon-Hyuk Chang
2024 A conf
INTERSPEECH
Jae-Hong Lee, Sang-Eon Lee, Dong-Hyun Kim, Do-Hee Kim, Joon-Hyuk Chang
2024 J jnl
IEEE ACM Trans. Audio Speech Lang. Process.
Jae-Hong Lee, Joon-Hyuk Chang
2024 A* conf
ICML
Jae-Hong Lee, Joon-Hyuk Chang
2024 Misc conf
ICASSP
Dong-Hyun Kim, Jae-Hong Lee, Joon-Hyuk Chang
2024 A conf
INTERSPEECH
Ji-Hun Kang, Jae-Hong Lee, Mun-Hak Lee, Joon-Hyuk Chang
2023 C conf
ASRU
Jae-Hong Lee, Do-Hee Kim, Joon-Hyuk Chang
2023 Misc conf
ICASSP
Jin-Seong Choi, Jae-Hong Lee, Chae-Won Lee, Joon-Hyuk Chang
2023 Misc conf
ICASSP
Jae-Hong Lee, Dong-Hyun Kim, Joon-Hyuk Chang
2022 A conf
INTERSPEECH
Jae-Hong Lee, Chae Won Lee, Jin-Seong Choi, Joon-Hyuk Chang, Woo Kyeong Seong, Jeonghan Lee
2022 J jnl
NeuroImage
Peter R. Millar, Patrick H. Luckett, Brian A. Gordon, Tammie L. S. Benzinger, Suzanne E. Schindler, Anne M. Fagan, Carlos Cruchaga, Randall J. Bateman, Ricardo Allegri, Mathias Jucker, Jae-Hong Lee, Hiroshi Mori, Stephen P. Salloway, Igor Yakushev, John C. Morris, Beau M. Ances
2022 A conf
INTERSPEECH
Dong-Hyun Kim, Jae-Hong Lee, Ji-Hwan Mo, Joon-Hyuk Chang
2021 J jnl
CoRR
Jae-Hong Lee, Joon-Hyuk Chang
2019 J jnl
BMC Medical Informatics Decis. Mak.
Min Ju Kang, Sang Yun Kim, Duk L. Na, Byeong C. Kim, Dong Won Yang, Eun-Joo Kim, Hae Ri Na, Hyun Jeong Han, Jae-Hong Lee, Jong Hun Kim, Kee Hyung Park, Kyung Won Park, Seol-Heui Han, Seong Yoon Kim, Soo Jin Yoon, Bora Yoon, Sang Won Seo, So Young Moon, Young-soon Yang, Yong S. Shim, Min Jae Baek, Jee Hyang Jeong, Seong Hye Choi, Young Chul Youn
2014 B conf
SMC
Hyoungrae Kim, Jae-Hong Lee, Hakil Kim, Daehyuk Park
2013 B conf
SMC
Jae-Hong Lee, Xuenan Cui, Seung-Jun Lee, Hakil Kim, Hyoungrae Kim
2013 conf
URAI
Hyoungrae Kim, Jae-Hong Lee, Seung-Jun Lee, Xuenan Cui, Hakil Kim
1997 conf
Integrated Network Management
Jong-Tae Park, Jae-Hong Lee, James Won-Ki Hong, Young-Myung Kim, Sung-Bum Kim
docs/CODE_ANALYSIS_APPROACH.md
← Index docs/CODE_ANALYSIS_APPROACH.md markdown
# Code Analysis Approach

This document explains the code analysis methodologies used in the REDB malware analysis framework.

## Disassembly Normalization

The framework implements a sophisticated three-level normalization strategy for disassembled code that provides different levels of abstraction for similarity detection and feature extraction.

### Overall Normalization Strategy

The framework implements a **hierarchical abstraction approach** where each instruction is normalized at three different levels simultaneously:

1. **Level 0 (fully_normalized)**: Maximum abstraction - reduces operands to broad categories
2. **Level 1 (api_normalized)**: Medium abstraction - preserves semantic meaning while normalizing details  
3. **Level 2 (category_normalized)**: Minimum abstraction - maintains architectural specificity

This multi-level approach allows analysts to perform similarity analysis at different granularities depending on their specific detection goals.

### Implementation Architecture

The normalization process follows this workflow:

1. **Token Parsing**: Each instruction is parsed from Binary Ninja's instruction tokens to extract the mnemonic and operands
2. **Multi-Level Processing**: Each operand is processed through all three normalization functions
3. **Instruction Reconstruction**: Normalized instructions are rebuilt with the mnemonic plus normalized operands
4. **Control Flow Tagging**: Control flow instructions get a `<TARGET>` suffix for easier pattern matching

### Level 0: Fully Normalized (Maximum Abstraction)

**Purpose**: Creates the most abstract representation for broad pattern detection across different malware families.

**Transformations**:
- **Registers**: All registers normalized to semantic categories via `normalize_register()`:
  - General purpose registers (EAX, EBX, R8, etc.) → `GPR`
  - Stack/Base pointers (ESP, EBP, RSP) → `PTR` 
  - SIMD registers (XMM0, XMM1) → `XMM`
  - FPU registers (ST0, ST1) → `FPU`
- **Memory Operations**: All memory references → `MEM`
- **Constants**: All immediate values → `CONST`  
- **Data References**: All symbols/data references → `DATA_REF`

**Example**:
```
mov eax, [ebp+8]     → MOV GPR MEM
call CreateFileW     → CALL DATA_REF <TARGET>
add ecx, 0x10        → ADD GPR CONST
```

### Level 1: API Normalized (Medium Abstraction)

**Purpose**: Preserves semantic distinctions while normalizing architectural details. Focuses on behavioral patterns and API usage.

**Transformations**:
- **Registers**: Categorized by functional role:
  - Data registers → `GPR_DATA`
  - Index registers (ESI, EDI) → `GPR_INDEX`  
  - Stack registers (ESP, EBP) → `GPR_STACK`
  - SIMD registers → `XMM_REG`
- **Memory Operations**: Classified by access pattern:
  - Stack access → `MEM_STACK`
  - String operations → `MEM_STRING` 
  - General access → `MEM_GENERAL`
- **Constants**: Categorized by range:
  - Small constants (-16 to 16) → `CONST_{value}`
  - Large constants → `CONST_LARGE`
- **API Calls**: Resolved to specific API names:
  - `CreateFileW` → `API_CreateFileW`
  - Other symbols → `DATA_SYM`

**Example**:
```
mov eax, [ebp+8]     → MOV GPR_DATA MEM_STACK
call CreateFileW     → CALL API_CreateFileW <TARGET>
add ecx, 0x10        → ADD GPR_DATA CONST_LARGE
```

### Level 2: Category Normalized (Minimum Abstraction)

**Purpose**: Maintains architectural specificity while normalizing specific values. Best for detecting variants with similar implementation details.

**Transformations**:
- **Registers**: Architecture-specific categories:
  - 64-bit registers → `REG_64`, with special cases for `REG_64_SP`, `REG_64_BP`
  - 32-bit registers → `REG_32`
  - 16/8-bit registers → `REG_16_8`
- **Memory Operations**: Detailed addressing mode classification:
  - Complex addressing → `MEM_SCALED_INDEX`
  - Base + offset → `MEM_BASE_OFFSET`
  - Direct addressing → `MEM_DIRECT`
- **Constants**: Type-specific classification:
  - Hexadecimal → `CONST_HEX`
  - Decimal → `CONST_DEC`
- **API Calls**: Categorized by functional group:
  - File operations → `API_FILE_OP`
  - Memory operations → `API_MEMORY_OP`
  - Network operations → `API_NETWORK_OP`

**Example**:
```
mov eax, [ebp+8]     → MOV REG_32 MEM_BASE_OFFSET
call CreateFileW     → CALL API_FILE_OP <TARGET>
add ecx, 0x10        → ADD REG_32 CONST_HEX
```

### Key Features and Benefits

#### 1. Multi-Granularity Similarity Detection
- **Level 0**: Detects broad behavioral patterns across malware families
- **Level 1**: Identifies API usage patterns and semantic similarities
- **Level 2**: Finds variants with similar implementation approaches

#### 2. Robust Pattern Matching
- Control flow instructions tagged with `<TARGET>` for easier CFG analysis
- Handles edge cases with fallback mechanisms
- Consistent uppercase normalization prevents case sensitivity issues

#### 3. API-Aware Analysis
The framework includes sophisticated API recognition through the `ApiCategory` enum and resolution methods:
- **File Operations**: CreateFile, ReadFile, WriteFile, etc.
- **Memory Operations**: VirtualAlloc, HeapAlloc, VirtualProtect, etc.  
- **Registry Operations**: RegOpenKey, RegSetValue, etc.
- **Network Operations**: WSASocket, send, recv, etc.
- **Process Operations**: CreateProcess, OpenProcess, etc.

#### 4. Scalable Feature Extraction
Each level produces different hash values for the same function:
- `fully_normalized_disassembly_hash`
- `api_normalized_disassembly_hash`  
- `category_normalized_disassembly_hash`

This enables efficient similarity searches at different abstraction levels in the ClickHouse database.

### Practical Applications for Malware Analysis

#### Threat Hunting Scenarios:

1. **Family Detection** (Level 0): Find samples using similar algorithmic approaches regardless of specific implementation
2. **Variant Analysis** (Level 1): Identify samples with similar API usage patterns and behavioral semantics
3. **Code Reuse Detection** (Level 2): Discover samples sharing specific implementation techniques or code fragments

#### Similarity Metrics Integration:
- Each normalization level can be used with different fuzzy hashing algorithms (ssdeep, TLSH, etc.)
- Level 0 works well with structural similarity metrics
- Level 1 optimal for behavioral similarity analysis  
- Level 2 suitable for implementation-specific pattern matching

This three-tiered approach provides malware analysts with flexible tools for detecting similarities across the threat landscape while maintaining the precision needed for detailed variant analysis.



---

*More code analysis approaches will be documented in additional sections as they are implemented.*