Hang Huang

19 papers A* 1A 5C 2Misc 1Journal 8Unranked 2
YearRankTypeTitle / Venue / Authors
2026 A conf
EuroSys
Zixuan Wang, Qi Wu, Hang Huang, Jia Rao, Hui Lu, Hao Fan, Zhuo Huang, Song Wu, Hai Jin
2025 J jnl
Future Gener. Comput. Syst.
Lichuan Ma, Lu Zhou, Hang Huang, Youyang Qu, Xuefeng Liu
2024 J jnl
IEEE Trans. Computers
Hang Huang, Honglei Wang, Jia Rao, Song Wu, Hao Fan, Chen Yu, Hai Jin, Kun Suo, Lisong Pan
2023 J jnl
IEEE Trans. Computers
Hang Huang, Yuqing Zhao, Jia Rao, Song Wu, Hai Jin, Duoqiang Wang, Kun Suo, Lisong Pan
2023 J jnl
Found. Comput. Math.
Austin Conner, Hang Huang, J. M. Landsberg
2023 J jnl
Future Gener. Comput. Syst.
Kun Wang, Song Wu, Kun Suo, Yijie Liu, Hang Huang, Zhuo Huang, Hai Jin
2023 J jnl
CoRR
Hang Huang, J. M. Landsberg
2023 A* conf
SOSP
Hang Huang, Jiangshan Lai, Jia Rao, Hui Lu, Wenlong Hou, Hang Su, Quan Xu, Jiang Zhong, Jiahao Zeng, Xu Wang, Zhengyu He, Weidong Han, Jiang Liu, Tao Ma, Song Wu
2022 Misc conf
TRIDENTCOM
Yafeng Li, Hang Huang, Lichuan Ma
2021 J jnl
CoRR
Haoran Zhou, Hang Huang, Rui Zhao, Wei Wang, Qingguo Zhou
2021 A conf
Middleware
Xiaofeng Wu, Jia Rao, Wei Chen, Hang Huang, Chris H. Q. Ding, Heng Huang
2021 A conf
HPDC
Hang Huang, Jia Rao, Song Wu, Hai Jin, Hong Jiang, Hao Che, Xiaofeng Wu
2020 conf
ISVC (2)
Hang Huang, Peng Zhi, Haoran Zhou, Yujin Zhang, Qiang Wu, Binbin Yong, Weijun Tan, Qingguo Zhou
2020 J jnl
CoRR
Austin Conner, Hang Huang, J. M. Landsberg
2019 A conf
HPDC
Hang Huang, Jia Rao, Song Wu, Hai Jin, Kun Suo, Xiaofeng Wu
2018 C conf
ISMM
Rodrigo Bruno, Paulo Ferreira, Ruslan Synytsky, Tetiana Fydorenchyk, Jia Rao, Hang Huang, Song Wu
2016 C conf
APSCC
Bowen Ruan, Hang Huang, Song Wu, Hai Jin
2007 conf
CICC
Daniel Murray, James Burnette, Brian Campbell, Mark Chung, Bruce Fernandes, Subhendra Ghosh, Rajat Goel, Greg Hess, Hang Huang, Zhibin Huang, Naveen Javarappa, Pradeep Kanapathipillai, Fabian Klass, Fang Liu, Anup Mehta, Yamini Modukuru, Nishant Nerurkar, Abhijit Radhakrishnan, Sribalan Santhanam, Junji Sugisawa, Shyam Sundar, Honkai John Tam, Ricky Wen, Eric Wu, Jung-Cheng Yeh, John Yong, Sanjay Zambare
2002 A conf
ITC
Shahin Nazarian, Hang Huang, Suriyaprakash Natarajan, Sandeep K. Gupta, Melvin A. Breuer
CLAUDE.md
← Index CLAUDE.md markdown
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## Project Overview

REDB (RationalEdge Samples DB) is a malware analysis framework that extracts features from PE (Portable Executable) files and stores them in ClickHouse database for analysis. It provides a comprehensive set of extractors for analyzing binary samples including PE headers, imports, resources, signatures, and decompiled code.

## Common Commands

### Development Setup
```bash
source venv/bin/activate

# Install dependencies
pip install -r requirements.txt

# Run the main application
python start.py --path /path/to/samples --repo sample_repo --index_prefix redb
```

### Analysis Commands
```bash
# Process a single file
python start.py --path /path/to/binary --repo test --index_prefix redb

# Process from S3 storage
python start.py --s3 --repo malpedia --index_prefix redb

# Process from S3 storage but only a subset of a specific repository
python start.py --s3 --repo "vx-itw" --s3-notes "ITW.0138" --index_prefix redb

# Run only decompilation
python start.py --path /path/to/binary --repo test --index_prefix redb --decompile

# Run specific modules
python start.py --path /path/to/binary --repo test --index_prefix redb --modules "BasicPropertiesExtractor,PEFeaturesExtractor"

# Run as Nomad job (for containerized deployment)
python start.py --nomad-job
```

### Testing
There are no formal unit tests. Testing is done by running the extractors on sample files in the `test_files/` directory.

## Architecture Overview

### Core Components

1. **Ingestor (`redb/ingestor.py`)**: Main orchestrator that handles file processing, multiprocessing, and coordinates extractors
2. **Extractors (`redb/extractors/`)**: Modular analysis components that extract specific features
3. **Database Exporters (`redb/extractors/database_exporters.py`)**: Handle data export to ClickHouse
4. **Settings (`redb/settings/`)**: Configuration management for database connections

### Extractor Architecture

All extractors inherit from the base `Extractor` class and implement:
- `extract()`: Main analysis logic
- `prepare_export_data()`: Format data for database export
- `get_clickhouse_table()`: Return target table name

Available extractors:
- **General**: BasicPropertiesExtractor, HashExtractor, DIEExtractor, CAPAExtractor
- **PE-specific**: PEFeaturesExtractor, PEImportExtractor, PEResourceExtractor, PEOverlayExtractor, PESectionExtractor, PESignatureExtractor, PEDotNetExtractor, PEInconstistencyTestsExtractor, PEExtraFindings
- **ELF**: ELFFeaturesExtractor, ELFSegmentExtractor, ELFSectionExtractor, ELFDependencyExtractor, ELFSymbolExtractor, ELFImportExtractor, ELFExportExtractor, ELFRelocationExtractor, ELFNotesExtractor
- **Mach-O**: MachOFeaturesExtractor, MachOSegmentExtractor, MachOImportExtractor, MachOExportExtractor, MachODylibExtractor, MachOSignatureExtractor
- **APK**: APKFeaturesExtractor, APKManifestExtractor, APKPermissionsExtractor, APKSignatureExtractor, APKDexExtractor, APKResourceExtractor, APKNativeLibExtractor, APKInconsistencyTestsExtractor
- **Decompilation**: DecompileBinja, DecompileAPK

### Database Schema

The project uses a comprehensive ClickHouse schema defined in `redb/redb_schema.yml` with tables for:
- Basic properties (`redb_basic_properties`)
- PE features (`redb_pe_features`, `redb_pe_imports`, `redb_pe_sections`, etc.)
- Decompiled code (`code_binja_decompiled_functions_content`, `code_binja_decompiled_functions_references`)
- CAPA analysis (`redb_capa`, `redb_capa_capabilities`)

Full schema documentation is available in `docs/database_schema.md`.

### Processing Modes

1. **Analysis Mode**: Extracts features using selected modules
2. **Decompile Mode**: Uses Binary Ninja for code decompilation
3. **S3 Mode**: Fetches samples from S3 storage based on catalog queries
4. **Nomad Job Mode**: Processes single jobs using environment variables for containerized deployment

### Configuration

Environment variables are used for configuration:
- Database connection: `CLICKHOUSE_HOST`, `CLICKHOUSE_PORT`, `CLICKHOUSE_USER`, `CLICKHOUSE_PASSWORD`
- S3 storage: `S3_ENDPOINT`, `S3_ACCESS_KEY`, `S3_SECRET_KEY`
- Processing: `BATCH_SIZE`, `REDB_TIMEOUT`, `DECOMPILE_WORKER_TIMEOUT`
- Nomad jobs: `JOB_ID`, `S3_KEY`, `S3_BUCKET`, `WORKER_TYPE`, `CALLBACK_URL`, `ANALYSIS_MODULES`

## Important Implementation Details

### Multiprocessing
- Uses `spawn` method for multiprocessing to avoid memory issues
- Worker processes have timeout handlers to prevent hanging
- Supports both batch processing and streaming processing modes

### Memory Management
- Implements aggressive garbage collection between batches
- Monitors swap usage and restarts worker pools when needed
- Kills stuck processes automatically

### Error Handling
- Comprehensive logging with per-file context
- Graceful handling of corrupted or unsupported files
- Automatic retry logic for database operations

### Security Context
This is a defensive security tool for malware analysis. It processes potentially malicious files in a controlled environment to extract features for detection and analysis purposes.

## Development Notes

- The codebase is optimized for processing large batches of malware samples
- Extractors are designed to be modular and can be run individually or in combination
- Database schema supports both normalized and denormalized views for different query patterns
- S3 integration allows for scalable processing of large malware repositories