Rajesh Challa

23 papers B 3Journal 11Unranked 9
YearRankTypeTitle / Venue / Authors
2026 J jnl
IEEE Trans. Cogn. Commun. Netw.
Sardar Jaffar Ali, Syed M. Raza, Duc-Tai Le, Rajesh Challa, Min Young Chung, Ness Shroff, Hyunseung Choo
2025 J jnl
CoRR
Sardar Jaffar Ali, Syed M. Raza, Duc-Tai Le, Rajesh Challa, Min Young Chung, Ness B. Shroff, Hyunseung Choo
2025 J jnl
Comput. Commun.
Sardar Jaffar Ali, Syed M. Raza, Huigyu Yang, Duc-Tai Le, Rajesh Challa, Moonseong Kim, Hyunseung Choo
2024 J jnl
IEEE Access
Ganesh Chandrasekaran, Abhinay Kumar, Karthikeyan Subramaniam, N. Sunil, Rajesh Challa
2023 J jnl
IEEE Access
Ganesh Chandrasekaran, Grace Hanusha Talakala, Shailendra Rajpoot, Rajesh Challa
2023 J jnl
J. Parallel Distributed Comput.
Mwasinga Lusungu Josh, Duc-Tai Le, Syed M. Raza, Rajesh Challa, Moonseong Kim, Hyunseung Choo
2023 J jnl
IEEE Netw.
Bokkeun Kim, Syed M. Raza, Huigyu Yang, Rajesh Challa, Duc-Tai Le, Hyunjun Choi, Dongjin Lee, Moonseong Kim, Hyunseung Choo
2022 J jnl
IEEE Access
Ganesh Chandrasekaran, Boopathi Ramasamy, Pranay Dhondi, Pankaj Bhimrao Thorat, Rajesh Challa
2021 conf
CCNC
Sivanandham Sivasankar, Rajesh Challa
2021 conf
CCNC
Pankaj Thorat, Niraj Kumar Dubey, Kunal Khetan, Rajesh Challa
2020 J jnl
Wirel. Networks
Syed M. Raza, Pankaj Thorat, Rajesh Challa, Seil Jeon, Hyunseung Choo
2020 J jnl
IEEE Syst. J.
Rajesh Challa, Vyacheslav V. Zalyubovskiy, Syed M. Raza, Hyunseung Choo, Aloknath De
2019 conf
IMCOM
Rajesh Challa, Syed M. Raza, Hyunseung Choo, Siwon Kim
2019 conf
IMCOM
Jisoo Lee, Syed M. Raza, Rajesh Challa, Jaeyeop Jeong, Hyunseung Choo
2019 conf
IMCOM
Bora Kim, Syed M. Raza, Rajesh Challa, Jongkwon Jang, Hyunseung Choo
2018 conf
NFV-SDN
Rajesh Challa, Syed M. Raza, Seil Jeon, Hyunseung Choo, Pankaj Thorat
2018 conf
IMCOM
Rajesh Challa, Seil Jeon, Syed M. Raza, Pankaj Thorat, Hyunseung Choo
2017 J jnl
IEEE Access
Rajesh Challa, Seil Jeon, Dongsoo S. Kim, Hyunseung Choo
2017 conf
ICOIN
Syed M. Raza, Pankaj Thorat, Rajesh Challa, Hyunseung Choo
2016 B conf
NetSoft
Rajesh Challa, Yongseung Lee, Hyunseung Choo
2016 B conf
NetSoft
Pankaj Thorat, Rajesh Challa, Syed M. Raza, Dongsoo S. Kim, Hyunseung Choo
2016 conf
IMCOM
Rajesh Challa, Syed M. Raza, Sangyep Nam, Hyunseung Choo
2016 B conf
NetSoft
Syed M. Raza, Pankaj Thorat, Rajesh Challa, Hyunseung Choo, Dongsoo S. Kim
docs/CODE_ANALYSIS_APPROACH.md
← Index docs/CODE_ANALYSIS_APPROACH.md markdown
# Code Analysis Approach

This document explains the code analysis methodologies used in the REDB malware analysis framework.

## Disassembly Normalization

The framework implements a sophisticated three-level normalization strategy for disassembled code that provides different levels of abstraction for similarity detection and feature extraction.

### Overall Normalization Strategy

The framework implements a **hierarchical abstraction approach** where each instruction is normalized at three different levels simultaneously:

1. **Level 0 (fully_normalized)**: Maximum abstraction - reduces operands to broad categories
2. **Level 1 (api_normalized)**: Medium abstraction - preserves semantic meaning while normalizing details  
3. **Level 2 (category_normalized)**: Minimum abstraction - maintains architectural specificity

This multi-level approach allows analysts to perform similarity analysis at different granularities depending on their specific detection goals.

### Implementation Architecture

The normalization process follows this workflow:

1. **Token Parsing**: Each instruction is parsed from Binary Ninja's instruction tokens to extract the mnemonic and operands
2. **Multi-Level Processing**: Each operand is processed through all three normalization functions
3. **Instruction Reconstruction**: Normalized instructions are rebuilt with the mnemonic plus normalized operands
4. **Control Flow Tagging**: Control flow instructions get a `<TARGET>` suffix for easier pattern matching

### Level 0: Fully Normalized (Maximum Abstraction)

**Purpose**: Creates the most abstract representation for broad pattern detection across different malware families.

**Transformations**:
- **Registers**: All registers normalized to semantic categories via `normalize_register()`:
  - General purpose registers (EAX, EBX, R8, etc.) → `GPR`
  - Stack/Base pointers (ESP, EBP, RSP) → `PTR` 
  - SIMD registers (XMM0, XMM1) → `XMM`
  - FPU registers (ST0, ST1) → `FPU`
- **Memory Operations**: All memory references → `MEM`
- **Constants**: All immediate values → `CONST`  
- **Data References**: All symbols/data references → `DATA_REF`

**Example**:
```
mov eax, [ebp+8]     → MOV GPR MEM
call CreateFileW     → CALL DATA_REF <TARGET>
add ecx, 0x10        → ADD GPR CONST
```

### Level 1: API Normalized (Medium Abstraction)

**Purpose**: Preserves semantic distinctions while normalizing architectural details. Focuses on behavioral patterns and API usage.

**Transformations**:
- **Registers**: Categorized by functional role:
  - Data registers → `GPR_DATA`
  - Index registers (ESI, EDI) → `GPR_INDEX`  
  - Stack registers (ESP, EBP) → `GPR_STACK`
  - SIMD registers → `XMM_REG`
- **Memory Operations**: Classified by access pattern:
  - Stack access → `MEM_STACK`
  - String operations → `MEM_STRING` 
  - General access → `MEM_GENERAL`
- **Constants**: Categorized by range:
  - Small constants (-16 to 16) → `CONST_{value}`
  - Large constants → `CONST_LARGE`
- **API Calls**: Resolved to specific API names:
  - `CreateFileW` → `API_CreateFileW`
  - Other symbols → `DATA_SYM`

**Example**:
```
mov eax, [ebp+8]     → MOV GPR_DATA MEM_STACK
call CreateFileW     → CALL API_CreateFileW <TARGET>
add ecx, 0x10        → ADD GPR_DATA CONST_LARGE
```

### Level 2: Category Normalized (Minimum Abstraction)

**Purpose**: Maintains architectural specificity while normalizing specific values. Best for detecting variants with similar implementation details.

**Transformations**:
- **Registers**: Architecture-specific categories:
  - 64-bit registers → `REG_64`, with special cases for `REG_64_SP`, `REG_64_BP`
  - 32-bit registers → `REG_32`
  - 16/8-bit registers → `REG_16_8`
- **Memory Operations**: Detailed addressing mode classification:
  - Complex addressing → `MEM_SCALED_INDEX`
  - Base + offset → `MEM_BASE_OFFSET`
  - Direct addressing → `MEM_DIRECT`
- **Constants**: Type-specific classification:
  - Hexadecimal → `CONST_HEX`
  - Decimal → `CONST_DEC`
- **API Calls**: Categorized by functional group:
  - File operations → `API_FILE_OP`
  - Memory operations → `API_MEMORY_OP`
  - Network operations → `API_NETWORK_OP`

**Example**:
```
mov eax, [ebp+8]     → MOV REG_32 MEM_BASE_OFFSET
call CreateFileW     → CALL API_FILE_OP <TARGET>
add ecx, 0x10        → ADD REG_32 CONST_HEX
```

### Key Features and Benefits

#### 1. Multi-Granularity Similarity Detection
- **Level 0**: Detects broad behavioral patterns across malware families
- **Level 1**: Identifies API usage patterns and semantic similarities
- **Level 2**: Finds variants with similar implementation approaches

#### 2. Robust Pattern Matching
- Control flow instructions tagged with `<TARGET>` for easier CFG analysis
- Handles edge cases with fallback mechanisms
- Consistent uppercase normalization prevents case sensitivity issues

#### 3. API-Aware Analysis
The framework includes sophisticated API recognition through the `ApiCategory` enum and resolution methods:
- **File Operations**: CreateFile, ReadFile, WriteFile, etc.
- **Memory Operations**: VirtualAlloc, HeapAlloc, VirtualProtect, etc.  
- **Registry Operations**: RegOpenKey, RegSetValue, etc.
- **Network Operations**: WSASocket, send, recv, etc.
- **Process Operations**: CreateProcess, OpenProcess, etc.

#### 4. Scalable Feature Extraction
Each level produces different hash values for the same function:
- `fully_normalized_disassembly_hash`
- `api_normalized_disassembly_hash`  
- `category_normalized_disassembly_hash`

This enables efficient similarity searches at different abstraction levels in the ClickHouse database.

### Practical Applications for Malware Analysis

#### Threat Hunting Scenarios:

1. **Family Detection** (Level 0): Find samples using similar algorithmic approaches regardless of specific implementation
2. **Variant Analysis** (Level 1): Identify samples with similar API usage patterns and behavioral semantics
3. **Code Reuse Detection** (Level 2): Discover samples sharing specific implementation techniques or code fragments

#### Similarity Metrics Integration:
- Each normalization level can be used with different fuzzy hashing algorithms (ssdeep, TLSH, etc.)
- Level 0 works well with structural similarity metrics
- Level 1 optimal for behavioral similarity analysis  
- Level 2 suitable for implementation-specific pattern matching

This three-tiered approach provides malware analysts with flexible tools for detecting similarities across the threat landscape while maintaining the precision needed for detailed variant analysis.



---

*More code analysis approaches will be documented in additional sections as they are implemented.*