Can Lu

29 papers A* 1Misc 2Journal 18Unranked 8
YearRankTypeTitle / Venue / Authors
2026 J jnl
CoRR
Keke Tang, Xianheng Liu, Weilong Peng, Xiaofei Wang, Daizong Liu, Peican Zhu, Can Lu, Zhihong Tian
2025 J jnl
IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens.
Can Lu, Feng Wang, Zhen Wang, Nan Xu, Zhuhong You, De-Shuang Huang
2025 J jnl
npj Digit. Medicine
Can Lu, Shenwei Wan, Zhiyong Liu
2025 Misc conf
WISA
Luyao Gao, Sijin Wang, Yikai Zhang, Can Lu
2025 J jnl
Proc. ACM Manag. Data
Can Lu, Sijin Wang, Wenxuan Deng, Xinyu Li, Yikai Zhang, Jeffrey Xu Yu
2025 J jnl
Int. J. Appl. Earth Obs. Geoinformation
Can Lu, Hanqing Xu, Qian Yao, Qing Liu, Jeremy D. Bricker, Sebastiaan N. Jonkman, Jie Yin, Jun Wang
2024 J jnl
Symmetry
Jinshui Wang, Yao Xin, Can Lu, Chengjun Jia, Yiming Ding
2024 A* conf
WWW
Siyi Teng, Jiadong Xie, Fan Zhang, Can Lu, Juntao Fang, Kai Wang
2023 J jnl
VLDB J.
Wenfei Fan, Yuanhao Li, Muyang Liu, Can Lu
2023 conf
TFC
Jianguang Sun, Bo Zhang, Can Lu, Ranye Du, Runze Miao
2022 conf
SIGMOD Conference
Wenfei Fan, Yuanhao Li, Muyang Liu, Can Lu
2022 J jnl
IEEE Trans. Knowl. Data Eng.
Can Lu, Jeffrey Xu Yu, Hao Wei, Yikai Zhang
2021 conf
SIGMOD Conference
Can Lu, Jeffrey Xu Yu, Zhiwei Zhang, Hong Cheng
2021 conf
SIGMOD Conference
Wenfei Fan, Yuanhao Li, Muyang Liu, Can Lu
2019 conf
SecureComm (2)
Xuetao Wei, Can Lu, Fatma Rana Ozcan, Ting Chen, Boyang Wang, Di Wu, Qiang Tang
2019 J jnl
CoRR
Can Lu, Jeffrey Xu Yu, Zhiwei Zhang, Hong Cheng
2018 conf
ICMSSP
Fengsong Hu, Xiajie Quan, Can Lu
2018 J jnl
IEEE Trans. Knowl. Data Eng.
Can Lu, Jeffrey Xu Yu, Hao Wei
2018 J jnl
VLDB J.
Hao Wei, Jeffrey Xu Yu, Can Lu, Ruoming Jin
2018 conf
ICMSSP
Can Lu, Bo Ou, Xiajie Quan
2018 J jnl
IEEE Trans. Knowl. Data Eng.
Hao Wei, Jeffrey Xu Yu, Can Lu
2017 J jnl
Int. J. Electron. Commer.
Tingting Christina Zhang, Can Lu, Murat Kizildag
2017 J jnl
Proc. VLDB Endow.
Can Lu, Jeffrey Xu Yu, Hao Wei, Yikai Zhang
2016 J jnl
IEEE Trans. Knowl. Data Eng.
Can Lu, Jeffrey Xu Yu, Rong-Hua Li, Hao Wei
2016 conf
SIGMOD Conference
Hao Wei, Jeffrey Xu Yu, Can Lu, Xuemin Lin
2016 J jnl
J. Intell. Fuzzy Syst.
Wei Li, Can Lu, Shuai Liu
2015 J jnl
CoRR
Can Lu, Jeffrey Xu Yu, Rong-Hua Li, Hao Wei
2014 J jnl
Proc. VLDB Endow.
Hao Wei, Jeffrey Xu Yu, Can Lu, Ruoming Jin
2005 Misc conf
AMIA
Byungsuk Choi, Stan Drozdetski, Margrethe Hackett, Can Lu, Cari Rottenberg, Linda Yu, Dale A. Hunscher, Daniel J. Clauw
CLAUDE.md
← Index CLAUDE.md markdown
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## Project Overview

REDB (RationalEdge Samples DB) is a malware analysis framework that extracts features from PE (Portable Executable) files and stores them in ClickHouse database for analysis. It provides a comprehensive set of extractors for analyzing binary samples including PE headers, imports, resources, signatures, and decompiled code.

## Common Commands

### Development Setup
```bash
source venv/bin/activate

# Install dependencies
pip install -r requirements.txt

# Run the main application
python start.py --path /path/to/samples --repo sample_repo --index_prefix redb
```

### Analysis Commands
```bash
# Process a single file
python start.py --path /path/to/binary --repo test --index_prefix redb

# Process from S3 storage
python start.py --s3 --repo malpedia --index_prefix redb

# Process from S3 storage but only a subset of a specific repository
python start.py --s3 --repo "vx-itw" --s3-notes "ITW.0138" --index_prefix redb

# Run only decompilation
python start.py --path /path/to/binary --repo test --index_prefix redb --decompile

# Run specific modules
python start.py --path /path/to/binary --repo test --index_prefix redb --modules "BasicPropertiesExtractor,PEFeaturesExtractor"

# Run as Nomad job (for containerized deployment)
python start.py --nomad-job
```

### Testing
There are no formal unit tests. Testing is done by running the extractors on sample files in the `test_files/` directory.

## Architecture Overview

### Core Components

1. **Ingestor (`redb/ingestor.py`)**: Main orchestrator that handles file processing, multiprocessing, and coordinates extractors
2. **Extractors (`redb/extractors/`)**: Modular analysis components that extract specific features
3. **Database Exporters (`redb/extractors/database_exporters.py`)**: Handle data export to ClickHouse
4. **Settings (`redb/settings/`)**: Configuration management for database connections

### Extractor Architecture

All extractors inherit from the base `Extractor` class and implement:
- `extract()`: Main analysis logic
- `prepare_export_data()`: Format data for database export
- `get_clickhouse_table()`: Return target table name

Available extractors:
- **General**: BasicPropertiesExtractor, HashExtractor, DIEExtractor, CAPAExtractor
- **PE-specific**: PEFeaturesExtractor, PEImportExtractor, PEResourceExtractor, PEOverlayExtractor, PESectionExtractor, PESignatureExtractor, PEDotNetExtractor, PEInconstistencyTestsExtractor, PEExtraFindings
- **ELF**: ELFFeaturesExtractor, ELFSegmentExtractor, ELFSectionExtractor, ELFDependencyExtractor, ELFSymbolExtractor, ELFImportExtractor, ELFExportExtractor, ELFRelocationExtractor, ELFNotesExtractor
- **Mach-O**: MachOFeaturesExtractor, MachOSegmentExtractor, MachOImportExtractor, MachOExportExtractor, MachODylibExtractor, MachOSignatureExtractor
- **APK**: APKFeaturesExtractor, APKManifestExtractor, APKPermissionsExtractor, APKSignatureExtractor, APKDexExtractor, APKResourceExtractor, APKNativeLibExtractor, APKInconsistencyTestsExtractor
- **Decompilation**: DecompileBinja, DecompileAPK

### Database Schema

The project uses a comprehensive ClickHouse schema defined in `redb/redb_schema.yml` with tables for:
- Basic properties (`redb_basic_properties`)
- PE features (`redb_pe_features`, `redb_pe_imports`, `redb_pe_sections`, etc.)
- Decompiled code (`code_binja_decompiled_functions_content`, `code_binja_decompiled_functions_references`)
- CAPA analysis (`redb_capa`, `redb_capa_capabilities`)

Full schema documentation is available in `docs/database_schema.md`.

### Processing Modes

1. **Analysis Mode**: Extracts features using selected modules
2. **Decompile Mode**: Uses Binary Ninja for code decompilation
3. **S3 Mode**: Fetches samples from S3 storage based on catalog queries
4. **Nomad Job Mode**: Processes single jobs using environment variables for containerized deployment

### Configuration

Environment variables are used for configuration:
- Database connection: `CLICKHOUSE_HOST`, `CLICKHOUSE_PORT`, `CLICKHOUSE_USER`, `CLICKHOUSE_PASSWORD`
- S3 storage: `S3_ENDPOINT`, `S3_ACCESS_KEY`, `S3_SECRET_KEY`
- Processing: `BATCH_SIZE`, `REDB_TIMEOUT`, `DECOMPILE_WORKER_TIMEOUT`
- Nomad jobs: `JOB_ID`, `S3_KEY`, `S3_BUCKET`, `WORKER_TYPE`, `CALLBACK_URL`, `ANALYSIS_MODULES`

## Important Implementation Details

### Multiprocessing
- Uses `spawn` method for multiprocessing to avoid memory issues
- Worker processes have timeout handlers to prevent hanging
- Supports both batch processing and streaming processing modes

### Memory Management
- Implements aggressive garbage collection between batches
- Monitors swap usage and restarts worker pools when needed
- Kills stuck processes automatically

### Error Handling
- Comprehensive logging with per-file context
- Graceful handling of corrupted or unsupported files
- Automatic retry logic for database operations

### Security Context
This is a defensive security tool for malware analysis. It processes potentially malicious files in a controlled environment to extract features for detection and analysis purposes.

## Development Notes

- The codebase is optimized for processing large batches of malware samples
- Extractors are designed to be modular and can be run individually or in combination
- Database schema supports both normalized and denormalized views for different query patterns
- S3 integration allows for scalable processing of large malware repositories