Xiaofen Jia

16 papers Journal 16
YearRankTypeTitle / Venue / Authors
2026 J jnl
IEEE Trans. Comput. Soc. Syst.
Wenjun Zheng, Shoufei Han, Xiaofen Jia, Edmond Qi Wu, Weiping Ding
2025 J jnl
Comput. Biol. Medicine
Xiaofen Jia, Wenjie Wang, Mei Zhang, Baiting Zhao
2025 J jnl
Multim. Syst.
Baiting Zhao, Yingying Shang, Xiaofen Jia, Zhenhuan Liang, Rui Hu
2025 J jnl
Eng. Appl. Artif. Intell.
Yuran Chen, Baiting Zhao, Xiaofen Jia, Tianbing Ma
2025 J jnl
IET Image Process.
Rui Hu, Xiaofen Jia, Xiaolei Han, Zailiang Jiang, Baiting Zhao
2025 J jnl
Int. J. Imaging Syst. Technol.
Xiaofen Jia, Wenyang Wang, Zhenhuan Liang, Baiting Zhao, Mei Zhang, Cong Wang
2025 J jnl
Vis. Comput.
Xiaofen Jia, Junjun Liu, Baiting Zhao, Zhenhuan Liang
2024 J jnl
Discov. Comput.
Zhenhuan Liang, Xiaofen Jia, Xiaolei Han, Baiting Zhao, Zhu Feng
2023 J jnl
Neural Process. Lett.
Xiaofen Jia, Zhenhuan Liang, Yongcun Guo, Yourui Huang, Baiting Zhao
2022 J jnl
Neural Process. Lett.
Baiting Zhao, Xiao Dong, Yongcun Guo, Xiaofen Jia, Yourui Huang
2022 J jnl
Neural Process. Lett.
Xiaofen Jia, Jianqiao Li, Baiting Zhao, Yongcun Guo, Yourui Huang
2021 J jnl
IEEE Access
Xiaofen Jia, Shengjie Du, Yongcun Guo, Yourui Huang, Baiting Zhao
2020 J jnl
IEEE Access
Baiting Zhao, Rui Hu, Xiaofen Jia, Yongcun Guo
2020 J jnl
Signal Process. Image Commun.
Yongcun Guo, Xiaofen Jia, Baiting Zhao, Huarong Chai, Yourui Huang
2019 J jnl
IET Image Process.
Xiaofen Jia, Yongcun Guo, Baiting Zhao, Yourui Huang
2018 J jnl
J. Electronic Imaging
Xiaofen Jia, Huarong Chai, Yongcun Guo, Yourui Huang, Baiting Zhao
CLAUDE.md
← Index CLAUDE.md markdown
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## Project Overview

REDB (RationalEdge Samples DB) is a malware analysis framework that extracts features from PE (Portable Executable) files and stores them in ClickHouse database for analysis. It provides a comprehensive set of extractors for analyzing binary samples including PE headers, imports, resources, signatures, and decompiled code.

## Common Commands

### Development Setup
```bash
source venv/bin/activate

# Install dependencies
pip install -r requirements.txt

# Run the main application
python start.py --path /path/to/samples --repo sample_repo --index_prefix redb
```

### Analysis Commands
```bash
# Process a single file
python start.py --path /path/to/binary --repo test --index_prefix redb

# Process from S3 storage
python start.py --s3 --repo malpedia --index_prefix redb

# Process from S3 storage but only a subset of a specific repository
python start.py --s3 --repo "vx-itw" --s3-notes "ITW.0138" --index_prefix redb

# Run only decompilation
python start.py --path /path/to/binary --repo test --index_prefix redb --decompile

# Run specific modules
python start.py --path /path/to/binary --repo test --index_prefix redb --modules "BasicPropertiesExtractor,PEFeaturesExtractor"

# Run as Nomad job (for containerized deployment)
python start.py --nomad-job
```

### Testing
There are no formal unit tests. Testing is done by running the extractors on sample files in the `test_files/` directory.

## Architecture Overview

### Core Components

1. **Ingestor (`redb/ingestor.py`)**: Main orchestrator that handles file processing, multiprocessing, and coordinates extractors
2. **Extractors (`redb/extractors/`)**: Modular analysis components that extract specific features
3. **Database Exporters (`redb/extractors/database_exporters.py`)**: Handle data export to ClickHouse
4. **Settings (`redb/settings/`)**: Configuration management for database connections

### Extractor Architecture

All extractors inherit from the base `Extractor` class and implement:
- `extract()`: Main analysis logic
- `prepare_export_data()`: Format data for database export
- `get_clickhouse_table()`: Return target table name

Available extractors:
- **General**: BasicPropertiesExtractor, HashExtractor, DIEExtractor, CAPAExtractor
- **PE-specific**: PEFeaturesExtractor, PEImportExtractor, PEResourceExtractor, PEOverlayExtractor, PESectionExtractor, PESignatureExtractor, PEDotNetExtractor, PEInconstistencyTestsExtractor, PEExtraFindings
- **ELF**: ELFFeaturesExtractor, ELFSegmentExtractor, ELFSectionExtractor, ELFDependencyExtractor, ELFSymbolExtractor, ELFImportExtractor, ELFExportExtractor, ELFRelocationExtractor, ELFNotesExtractor
- **Mach-O**: MachOFeaturesExtractor, MachOSegmentExtractor, MachOImportExtractor, MachOExportExtractor, MachODylibExtractor, MachOSignatureExtractor
- **APK**: APKFeaturesExtractor, APKManifestExtractor, APKPermissionsExtractor, APKSignatureExtractor, APKDexExtractor, APKResourceExtractor, APKNativeLibExtractor, APKInconsistencyTestsExtractor
- **Decompilation**: DecompileBinja, DecompileAPK

### Database Schema

The project uses a comprehensive ClickHouse schema defined in `redb/redb_schema.yml` with tables for:
- Basic properties (`redb_basic_properties`)
- PE features (`redb_pe_features`, `redb_pe_imports`, `redb_pe_sections`, etc.)
- Decompiled code (`code_binja_decompiled_functions_content`, `code_binja_decompiled_functions_references`)
- CAPA analysis (`redb_capa`, `redb_capa_capabilities`)

Full schema documentation is available in `docs/database_schema.md`.

### Processing Modes

1. **Analysis Mode**: Extracts features using selected modules
2. **Decompile Mode**: Uses Binary Ninja for code decompilation
3. **S3 Mode**: Fetches samples from S3 storage based on catalog queries
4. **Nomad Job Mode**: Processes single jobs using environment variables for containerized deployment

### Configuration

Environment variables are used for configuration:
- Database connection: `CLICKHOUSE_HOST`, `CLICKHOUSE_PORT`, `CLICKHOUSE_USER`, `CLICKHOUSE_PASSWORD`
- S3 storage: `S3_ENDPOINT`, `S3_ACCESS_KEY`, `S3_SECRET_KEY`
- Processing: `BATCH_SIZE`, `REDB_TIMEOUT`, `DECOMPILE_WORKER_TIMEOUT`
- Nomad jobs: `JOB_ID`, `S3_KEY`, `S3_BUCKET`, `WORKER_TYPE`, `CALLBACK_URL`, `ANALYSIS_MODULES`

## Important Implementation Details

### Multiprocessing
- Uses `spawn` method for multiprocessing to avoid memory issues
- Worker processes have timeout handlers to prevent hanging
- Supports both batch processing and streaming processing modes

### Memory Management
- Implements aggressive garbage collection between batches
- Monitors swap usage and restarts worker pools when needed
- Kills stuck processes automatically

### Error Handling
- Comprehensive logging with per-file context
- Graceful handling of corrupted or unsupported files
- Automatic retry logic for database operations

### Security Context
This is a defensive security tool for malware analysis. It processes potentially malicious files in a controlled environment to extract features for detection and analysis purposes.

## Development Notes

- The codebase is optimized for processing large batches of malware samples
- Extractors are designed to be modular and can be run individually or in combination
- Database schema supports both normalized and denormalized views for different query patterns
- S3 integration allows for scalable processing of large malware repositories