Wei-ming Wang

18 papers B 1C 2Misc 1Journal 12Unranked 2
YearRankTypeTitle / Venue / Authors
2024 J jnl
Commun. Stat. Simul. Comput.
Wei-ming Wang, Chun-Che Wen, Rajdeep Das, Miin-Jye Wen
2024 J jnl
Commun. Stat. Simul. Comput.
Miin-Jye Wen, Chun-Che Wen, Wei-ming Wang
2021 J jnl
Soft Comput.
Xin Huang, Hongzhuan Chen, Peng Ma, Wei-ming Wang, Xiang Cai, Malik Nafis
2016 J jnl
Frontiers Inf. Technol. Electron. Eng.
Qingfeng Li, Shao-bo Chen, Wei-ming Wang, Hong-Wei Hao, Luming Li
2014 J jnl
Expert Syst. Appl.
Wei-ming Wang, Xun Peng, Guo-Niu Zhu, Jie Hu, Ying-hong Peng
2013 J jnl
Decis. Support Syst.
Wei-ming Wang, Amy H. I. Lee, Li-Pei Peng, Zih-Ling Wu
2012 conf
AusCTW
Wei-Lung Mao, Wei-ming Wang, Jyh Sheen, Po-Hung Chen
2011 J jnl
Expert Syst. Appl.
Jin Qi, Jie Hu, Ying-hong Peng, Wei-ming Wang, Zhenfei Zhang
2011 J jnl
Expert Syst. Appl.
Jin Qi, Jie Hu, Ying-hong Peng, Qiushi Ren, Wei-ming Wang, Zhenfei Zhang
2010 J jnl
Optim. Lett.
Wei-ming Wang, Amy H. I. Lee, Ding-Tsair Chang
2010 J jnl
Eur. J. Oper. Res.
Chung-Chi Hsieh, Yu-Te Liu, Wei-ming Wang
2009 J jnl
Expert Syst. Appl.
Jin Qi, Jie Hu, Ying-hong Peng, Wei-ming Wang, Zhenfei Zhang
2008 J jnl
IEEE Trans. Inf. Theory
Xiao-yu Chen, Wei-ming Wang
2007 conf
IMECS
Amy H. I. Lee, Tasi-Ying Lin, Wen-Chin Chen, Wei-ming Wang
2006 C conf
CSCWD
Jie Hu, Guangleng Xiong, Ying-hong Peng, Wei-ming Wang
2006 C conf
CSCWD
Wei-ming Wang, Jie Hu, Jilong Yin, Ying-hong Peng
2005 B conf
SMC
Wei-Yen Wang, I-Hsum Li, Wei-ming Wang, Shun-Feng Su, Nai-Jian Wang
2005 Misc conf
ICMLC
Wei-ming Wang, Jie Hu, Fei Zhou, Dayong Li, Xiangjun Fu, Ying-hong Peng
CLAUDE.md
← Index CLAUDE.md markdown
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## Project Overview

REDB (RationalEdge Samples DB) is a malware analysis framework that extracts features from PE (Portable Executable) files and stores them in ClickHouse database for analysis. It provides a comprehensive set of extractors for analyzing binary samples including PE headers, imports, resources, signatures, and decompiled code.

## Common Commands

### Development Setup
```bash
source venv/bin/activate

# Install dependencies
pip install -r requirements.txt

# Run the main application
python start.py --path /path/to/samples --repo sample_repo --index_prefix redb
```

### Analysis Commands
```bash
# Process a single file
python start.py --path /path/to/binary --repo test --index_prefix redb

# Process from S3 storage
python start.py --s3 --repo malpedia --index_prefix redb

# Process from S3 storage but only a subset of a specific repository
python start.py --s3 --repo "vx-itw" --s3-notes "ITW.0138" --index_prefix redb

# Run only decompilation
python start.py --path /path/to/binary --repo test --index_prefix redb --decompile

# Run specific modules
python start.py --path /path/to/binary --repo test --index_prefix redb --modules "BasicPropertiesExtractor,PEFeaturesExtractor"

# Run as Nomad job (for containerized deployment)
python start.py --nomad-job
```

### Testing
There are no formal unit tests. Testing is done by running the extractors on sample files in the `test_files/` directory.

## Architecture Overview

### Core Components

1. **Ingestor (`redb/ingestor.py`)**: Main orchestrator that handles file processing, multiprocessing, and coordinates extractors
2. **Extractors (`redb/extractors/`)**: Modular analysis components that extract specific features
3. **Database Exporters (`redb/extractors/database_exporters.py`)**: Handle data export to ClickHouse
4. **Settings (`redb/settings/`)**: Configuration management for database connections

### Extractor Architecture

All extractors inherit from the base `Extractor` class and implement:
- `extract()`: Main analysis logic
- `prepare_export_data()`: Format data for database export
- `get_clickhouse_table()`: Return target table name

Available extractors:
- **General**: BasicPropertiesExtractor, HashExtractor, DIEExtractor, CAPAExtractor
- **PE-specific**: PEFeaturesExtractor, PEImportExtractor, PEResourceExtractor, PEOverlayExtractor, PESectionExtractor, PESignatureExtractor, PEDotNetExtractor, PEInconstistencyTestsExtractor, PEExtraFindings
- **ELF**: ELFFeaturesExtractor, ELFSegmentExtractor, ELFSectionExtractor, ELFDependencyExtractor, ELFSymbolExtractor, ELFImportExtractor, ELFExportExtractor, ELFRelocationExtractor, ELFNotesExtractor
- **Mach-O**: MachOFeaturesExtractor, MachOSegmentExtractor, MachOImportExtractor, MachOExportExtractor, MachODylibExtractor, MachOSignatureExtractor
- **APK**: APKFeaturesExtractor, APKManifestExtractor, APKPermissionsExtractor, APKSignatureExtractor, APKDexExtractor, APKResourceExtractor, APKNativeLibExtractor, APKInconsistencyTestsExtractor
- **Decompilation**: DecompileBinja, DecompileAPK

### Database Schema

The project uses a comprehensive ClickHouse schema defined in `redb/redb_schema.yml` with tables for:
- Basic properties (`redb_basic_properties`)
- PE features (`redb_pe_features`, `redb_pe_imports`, `redb_pe_sections`, etc.)
- Decompiled code (`code_binja_decompiled_functions_content`, `code_binja_decompiled_functions_references`)
- CAPA analysis (`redb_capa`, `redb_capa_capabilities`)

Full schema documentation is available in `docs/database_schema.md`.

### Processing Modes

1. **Analysis Mode**: Extracts features using selected modules
2. **Decompile Mode**: Uses Binary Ninja for code decompilation
3. **S3 Mode**: Fetches samples from S3 storage based on catalog queries
4. **Nomad Job Mode**: Processes single jobs using environment variables for containerized deployment

### Configuration

Environment variables are used for configuration:
- Database connection: `CLICKHOUSE_HOST`, `CLICKHOUSE_PORT`, `CLICKHOUSE_USER`, `CLICKHOUSE_PASSWORD`
- S3 storage: `S3_ENDPOINT`, `S3_ACCESS_KEY`, `S3_SECRET_KEY`
- Processing: `BATCH_SIZE`, `REDB_TIMEOUT`, `DECOMPILE_WORKER_TIMEOUT`
- Nomad jobs: `JOB_ID`, `S3_KEY`, `S3_BUCKET`, `WORKER_TYPE`, `CALLBACK_URL`, `ANALYSIS_MODULES`

## Important Implementation Details

### Multiprocessing
- Uses `spawn` method for multiprocessing to avoid memory issues
- Worker processes have timeout handlers to prevent hanging
- Supports both batch processing and streaming processing modes

### Memory Management
- Implements aggressive garbage collection between batches
- Monitors swap usage and restarts worker pools when needed
- Kills stuck processes automatically

### Error Handling
- Comprehensive logging with per-file context
- Graceful handling of corrupted or unsupported files
- Automatic retry logic for database operations

### Security Context
This is a defensive security tool for malware analysis. It processes potentially malicious files in a controlled environment to extract features for detection and analysis purposes.

## Development Notes

- The codebase is optimized for processing large batches of malware samples
- Extractors are designed to be modular and can be run individually or in combination
- Database schema supports both normalized and denormalized views for different query patterns
- S3 integration allows for scalable processing of large malware repositories