Haibo Ni

18 papers Journal 4Unranked 14
YearRankTypeTitle / Venue / Authors
2026 J jnl
Briefings Bioinform.
Huasen Jiang, Xiaoyu Huang, Xiangpeng Bi, Wenjian Ma, Haibo Ni, Zhiqiang Wei, Pin Sun, Henggui Zhang, Shugang Zhang
2018 conf
CinC
Erick Andres Perez-Alday, Haibo Ni, Christopher Hamilton, Annabel Li-Pershing, Bernard Jaar, Jose Manuel Monroy Trujillo, Michelle Estrella, Rulan Parekh, Henggui Zhang, Larisa G. Tereshchenko
2018 conf
CinC
Inas A. Al Nemi, Haibo Ni, Henggui Zhang
2018 J jnl
PLoS Comput. Biol.
Wei Wang, Shanzhuo Zhang, Haibo Ni, Clifford J. Garratt, Mark R. Boyett, Jules C. Hancox, Henggui Zhang
2017 J jnl
PLoS Comput. Biol.
Dominic G. Whittaker, Haibo Ni, Aziza El Harchi, Jules C. Hancox, Henggui Zhang
2017 J jnl
PLoS Comput. Biol.
Michael A. Colman, Haibo Ni, Bo Liang, Nicole Schmitt, Henggui Zhang
2016 conf
CinC
Dominic G. Whittaker, Haibo Ni, Alan P. Benson, Jules C. Hancox, Henggui Zhang
2016 conf
CinC
Haibo Ni, Dominic G. Whittaker, Wei Wang, Henggui Zhang
2016 conf
CinC
Wei Wang, Dominic G. Whittaker, Haibo Ni, Kuanquan Wang, Henggui Zhang
2015 conf
CCECE
Haibo Ni, Lei Wang, Zhenjie Yang, Ming Zhu, Feng Ding, Zhenquan Qin
2015 conf
CinC
Erick Andres Perez-Alday, Chen Zhang, Michael A. Colman, Haibo Ni, Zizhao Gan, Henggui Zhang
2015 conf
CinC
Dominic G. Whittaker, Michael A. Colman, Haibo Ni, Jules C. Hancox, Henggui Zhang
2015 conf
CinC
Haibo Ni, Wei Wang, Erick Andres Perez-Alday, Henggui Zhang
2015 conf
CinC
Wei Wang, Haibo Ni, Henggui Zhang
2014 conf
CinC
Sanjay Kharche, Edward J. Vigmond, Haibo Ni, Michael A. Colman, Henggui Zhang
2014 conf
CinC
Haibo Ni, Michael A. Colman, Henggui Zhang
2013 conf
ICTC
Huaixian Yin, Haibo Ni, Liang Sun, Mingxuan Wang, Xun Zhou
2013 conf
CinC
Haibo Ni, Simon J. Castro, Robert S. Stephenson, Jonathan C. Jarvis, Tristan Lowe, George Hart, Mark R. Boyett, Henggui Zhang
Docker-README.md
← Index Docker-README.md markdown
# REDB Docker Setup

This document describes the Docker containerization for the REDB malware analysis framework.

## Overview

REDB has been containerized as a single unified image that supports both feature extraction and decompilation analysis. The container is stateless, processes files from S3 or local mounts, and exports results to ClickHouse database or via API callbacks.

## Architecture

- **Single Unified Container**: One image handles both feature extraction and decompilation
- **Runtime Tool Installation**: Tools (CAPA, DIE, Binary Ninja) installed at runtime from host snapshots
- **Stateless Processing**: No persistent storage required between runs
- **Multiple Invocation Modes**: Supports `--nomad-job`, `--s3`, `--s3-solo`, and `--path` modes
- **External Dependencies**: Connects to external ClickHouse and S3 services

## Files Structure

```
├── Dockerfile                 # Single unified container definition
├── docker-build.sh            # Build script with Docker Desktop bug workaround
├── docker-push.sh             # Push script to registry
├── test-docker.sh             # Container testing script
├── test-nomad.sh              # Nomad job mode testing
├── .dockerignore              # Build context exclusions
└── scripts/
    └── setup-and-run.sh       # Runtime tool setup entrypoint
```

## Tool Installation Strategy

The container uses a **runtime installation** approach:

1. **Base Image**: Contains Python dependencies and REDB code
2. **Runtime Setup**: `scripts/setup-and-run.sh` configures tools at container start
3. **Host Snapshots**: Binary Ninja installed from `/opt/binaryninja` if available
4. **System Tools**: CAPA and DIE expected at `/usr/bin/capa` and `/usr/bin/nfdc`

## Build and Run

### 1. Build Container

```bash
# Build unified image
./docker-build.sh

# Manual build
docker build --platform linux/amd64 -f Dockerfile -t redb:latest .
```

### 2. Run Modes

#### Nomad Job Mode (Primary)
```bash
# Feature extraction
docker run --rm \
  -e JOB_ID="analysis_001" \
  -e S3_KEY="samples/malware.exe" \
  -e S3_BUCKET="malware-bucket" \
  -e WORKER_TYPE="feature_extraction" \
  -e CALLBACK_URL="https://api.example.com/callbacks" \
  -e ANALYSIS_MODULES="BasicPropertiesExtractor,PEFeaturesExtractor" \
  -e CLICKHOUSE_HOST="clickhouse.example.com" \
  -e S3_ENDPOINT="s3.example.com" \
  -e S3_ACCESS_KEY="your-key" \
  -e S3_SECRET_KEY="your-secret" \
  redb:latest python3 start.py --nomad-job

# Decompilation (same container, different flags)
docker run --rm \
  -e JOB_ID="analysis_002" \
  -e S3_KEY="samples/malware.exe" \
  -e S3_BUCKET="malware-bucket" \
  -e WORKER_TYPE="decompilation" \
  -e CALLBACK_URL="https://api.example.com/callbacks" \
  -e ANALYSIS_MODULES="all" \
  -v /opt/binaryninja:/opt/binaryninja:ro \
  redb:latest python3 start.py --nomad-job --decompile
```

#### S3 Solo Mode
```bash
# Process single sample by S3 key (standard sharded path)
docker run --rm \
  -e S3_BUCKET="samples-bucket" \
  -e CLICKHOUSE_HOST="clickhouse.example.com" \
  -e S3_ENDPOINT="s3.example.com" \
  -e INDEX_PREFIX="redb" \
  -e REPO="test-analysis" \
  redb:latest python3 start.py --s3-solo "09/f7/09f7d02a3c2382199458c98a62b045145ee54ab6aba86166aecf3d10c3c1444c.zip"

# Process private sample (with prepath)
docker run --rm \
  -e S3_BUCKET="samples-bucket" \
  -e CLICKHOUSE_HOST="clickhouse.example.com" \
  -e S3_ENDPOINT="s3.example.com" \
  -e INDEX_PREFIX="redb" \
  -e REPO="test-analysis" \
  redb:latest python3 start.py --s3-solo "private/ab/cd/abcd1234567890abcdef1234567890abcdef1234567890abcdef123456.zip"
```

#### Local Files Mode
```bash
# Mount local samples
docker run --rm \
  -v /path/to/samples:/samples:ro \
  -v ./logs:/app/logs \
  -e CLICKHOUSE_HOST="clickhouse.example.com" \
  redb:latest python3 start.py --path /samples --repo local_test --index_prefix redb
```

## Environment Variables

### Required for Nomad Job Mode
- `JOB_ID` - Unique job identifier
- `S3_KEY` - S3 object key for sample
- `S3_BUCKET` - S3 bucket name
- `WORKER_TYPE` - "feature_extraction" or "decompilation"
- `CALLBACK_URL` - API endpoint for results
- `ANALYSIS_MODULES` - Comma-separated extractor list or "all"

### Database Configuration
- `CLICKHOUSE_HOST` - ClickHouse server hostname
- `CLICKHOUSE_PORT` - Port (default: 8123)
- `CLICKHOUSE_USER` - Database user (default: default)
- `CLICKHOUSE_PASSWORD` - Database password
- `CLICKHOUSE_DATABASE` - Database name (default: default)

### S3 Configuration
- `S3_ENDPOINT` - S3 endpoint URL
- `S3_ACCESS_KEY` - S3 access key
- `S3_SECRET_KEY` - S3 secret key
- `S3_SECURE` - "true" or "false" for HTTPS

### Processing Configuration
- `INDEX_PREFIX` - Database table prefix (default: redb)
- `REPO` - Repository identifier for this analysis batch
- `BATCH_SIZE` - Processing batch size (default: 10)
- `REDB_TIMEOUT` - Analysis timeout in seconds (default: 300)

### Tool Timeouts
- `CAPA_TIMEOUT` - CAPA analysis timeout (default: 300)
- `DIE_TIMEOUT` - DIE analysis timeout (default: 180)
- `BINJA_TIMEOUT` - Binary Ninja timeout (default: 1200)
- `DECOMPILE_EXTRACTOR_TIMEOUT` - Decompilation timeout (default: 2580)

## Binary Ninja Setup

For decompilation capabilities, mount Binary Ninja from host:

```bash
# Mount Binary Ninja installation
-v /opt/binaryninja:/opt/binaryninja:ro

# Mount license file
-v /path/to/license.dat:/home/analyzer/.binaryninja/license.dat:ro
```

The container will automatically detect and configure Binary Ninja at runtime.

## Registry Deployment

### Push to Registry
```bash
# Tag and push
./docker-push.sh

# Or manually
docker tag redb:latest your-registry/redb:latest
docker push your-registry/redb:latest
```

### Pull and Run
```bash
docker pull your-registry/redb:latest
docker run your-registry/redb:latest python3 start.py --nomad-job
```

## Testing

### Container Functionality Test
```bash
# Test with S3 key (standard sharded path)
./test-docker.sh "09/f7/09f7d02a3c2382199458c98a62b045145ee54ab6aba86166aecf3d10c3c1444c.zip"

# Test with private sample S3 key
./test-docker.sh "private/ab/cd/abcd1234567890abcdef1234567890abcdef1234567890abcdef123456.zip"
```

### Nomad Job Architecture Test
```bash
# Test Nomad job mode
./test-nomad.sh
```

## Development

### Interactive Container
```bash
# Debug container interactively
docker run -it --entrypoint /bin/bash redb:latest

# Check tool availability
docker run --rm redb:latest which python3
docker run --rm redb:latest ls -la /usr/bin/capa
```

### Build Troubleshooting

The build script includes workarounds for Docker Desktop bugs:

```bash
# If build hangs at "exporting to image", press Ctrl+C
# The image will still be created and tagged automatically
./docker-build.sh
```

### Container Logs
```bash
# View logs from mounted directory
docker run -v ./logs:/app/logs redb:latest python3 start.py --path /samples
tail -f logs/*.txt
```

## Production Notes

### Resource Requirements
- **Memory**: 2-4GB recommended (8GB for decompilation)
- **CPU**: 2+ cores recommended
- **Disk**: Minimal (stateless container)
- **Network**: Access to ClickHouse and S3 services

### Security
- Container runs as non-root user `analyzer` (UID 1000)
- Sample files should be mounted read-only
- No persistent state between container runs
- Isolated processing environment for malware analysis

### Deployment Architecture

This container is designed for:
- **Nomad job dispatch**: Single-use containers processing one sample each
- **Kubernetes jobs**: Batch processing with external orchestration
- **CI/CD pipelines**: Automated analysis in build systems
- **Development**: Local testing and debugging

The unified container approach means the same image handles both feature extraction and decompilation - the difference is only in the command-line flags used when starting the container.