Hannu Koivisto

30 papers B 1C 1Misc 1Journal 13Unranked 13
YearRankTypeTitle / Venue / Authors
2019 J jnl
Comput. Ind. Eng.
Idelfonso B. R. Nogueira, Márcio A. F. Martins, Reiner Requião, Amanda R. Oliveira, Vinícius Viena, Hannu Koivisto, Alírio E. Rodrigues, José M. Loureiro, Ana M. Ribeiro
2019 conf
ISGT Europe
Shengye Lu, Sami Repo, Mikko Salmenperä, Jari Seppälä, Hannu Koivisto
2018 J jnl
Appl. Soft Comput.
Idelfonso B. R. Nogueira, Ana M. Ribeiro, Reiner Requião, Karen V. Pontes, Hannu Koivisto, Alírio E. Rodrigues, José M. Loureiro
2018 J jnl
Comput. Chem. Eng.
Idelfonso B. R. Nogueira, Rui P. V. Faria, Reiner Requião, Hannu Koivisto, Márcio A. F. Martins, Alírio E. Rodrigues, José M. Loureiro, Ana M. Ribeiro
2017 conf
ICIT
Peyman Jafary, Ontrei Raipala, Sami Repo, Mikko Salmenperä, Jari Seppälä, Hannu Koivisto, Seppo Horsmanheimo, Heli Kokkoniemi-Tarkkanen, Lotta Tuomimäki, Amelia Alvarez de Sotomayor, Francisco Ramos, Alessio Dede, Davide Della Giustina
2017 conf
ICIT
Jari Seppälä, Hannu Koivisto, Peyman Jafary, Sami Repo
2016 conf
ICIT
Peyman Jafary, Sami Repo, Hannu Koivisto
2015 C conf
INDIN
Peyman Jafary, Sami Repo, Mikko Salmenperä, Hannu Koivisto
2015 conf
ISGT Asia
Peyman Jafary, Sami Repo, Hannu Koivisto
2011 J jnl
Appl. Soft Comput.
Tomi Helin, Hannu Koivisto
2010 J jnl
IEEE Trans. Fuzzy Syst.
Pietari Pulkkinen, Hannu Koivisto
2010 J jnl
Appl. Artif. Intell.
Antti Hahto, Jouni Mattila, Hannu Koivisto
2009 J jnl
Simul. Model. Pract. Theory
Tomi Roinila, Tomi Helin, Matti Vilkko, Teuvo Suntio, Hannu Koivisto
2009 conf
ICINCO-RA
Olli Post, Jari Seppälä, Hannu Koivisto
2008 J jnl
Eng. Appl. Artif. Intell.
Pietari Pulkkinen, Jarmo Hytönen, Hannu Koivisto
2008 J jnl
Int. J. Approx. Reason.
Pietari Pulkkinen, Hannu Koivisto
2008 conf
ITI
Mikko Laurikkala, Hannu Koivisto
2008 J jnl
Knowl. Based Syst.
Pietari Pulkkinen, Mikko Laurikkala, Aino Ropponen, Hannu Koivisto
2007 J jnl
Appl. Soft Comput.
Pietari Pulkkinen, Hannu Koivisto
2006 B conf
FUZZ-IEEE
Pietari Pulkkinen, Hannu Koivisto
2005 ch.
Classification and Clustering for Knowledge Discovery
Timo Lampinen, Mikko Laurikkala, Hannu Koivisto, Tapani Honkanen
2004 conf
ICINCO (1)
Henri Helanterä, Mikko Salmenperä, Hannu Koivisto
2004 conf
ICINCO (1)
Heikki Rasku, Juuso Rantala, Hannu Koivisto
2003 conf
MWCN
Eero Wallenius, Timo Hämäläinen, Timo Nihtilä, Jyrki Joutsensalo, Hannu Koivisto
2003 Misc conf
International Conference on Computational Science
Guangyu Xiong, Hannu Koivisto
2003 conf
ICECS
Eero Wallenius, Timo Hämäläinen, Jyrki Joutsensalo, Hannu Koivisto
2002 conf
FSKD
Timo Lampinen, Hannu Koivisto, Tapani Honkanen
2000 conf
ICC (1)
Heikki Vatiainen, Jarmo Harju, Hannu Koivisto, Sampo Saaristo, Juha Vihervaara
1999 J jnl
Intell. Autom. Soft Comput.
Ari S. Nissinen, Heikki N. Koivo, Hannu Koivisto
1991 J jnl
IEEE Trans. Syst. Man Cybern.
Timo Sorsa, Heikki N. Koivo, Hannu Koivisto
CLAUDE.md
← Index CLAUDE.md markdown
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## Project Overview

REDB (RationalEdge Samples DB) is a malware analysis framework that extracts features from PE (Portable Executable) files and stores them in ClickHouse database for analysis. It provides a comprehensive set of extractors for analyzing binary samples including PE headers, imports, resources, signatures, and decompiled code.

## Common Commands

### Development Setup
```bash
source venv/bin/activate

# Install dependencies
pip install -r requirements.txt

# Run the main application
python start.py --path /path/to/samples --repo sample_repo --index_prefix redb
```

### Analysis Commands
```bash
# Process a single file
python start.py --path /path/to/binary --repo test --index_prefix redb

# Process from S3 storage
python start.py --s3 --repo malpedia --index_prefix redb

# Process from S3 storage but only a subset of a specific repository
python start.py --s3 --repo "vx-itw" --s3-notes "ITW.0138" --index_prefix redb

# Run only decompilation
python start.py --path /path/to/binary --repo test --index_prefix redb --decompile

# Run specific modules
python start.py --path /path/to/binary --repo test --index_prefix redb --modules "BasicPropertiesExtractor,PEFeaturesExtractor"

# Run as Nomad job (for containerized deployment)
python start.py --nomad-job
```

### Testing
There are no formal unit tests. Testing is done by running the extractors on sample files in the `test_files/` directory.

## Architecture Overview

### Core Components

1. **Ingestor (`redb/ingestor.py`)**: Main orchestrator that handles file processing, multiprocessing, and coordinates extractors
2. **Extractors (`redb/extractors/`)**: Modular analysis components that extract specific features
3. **Database Exporters (`redb/extractors/database_exporters.py`)**: Handle data export to ClickHouse
4. **Settings (`redb/settings/`)**: Configuration management for database connections

### Extractor Architecture

All extractors inherit from the base `Extractor` class and implement:
- `extract()`: Main analysis logic
- `prepare_export_data()`: Format data for database export
- `get_clickhouse_table()`: Return target table name

Available extractors:
- **General**: BasicPropertiesExtractor, HashExtractor, DIEExtractor, CAPAExtractor
- **PE-specific**: PEFeaturesExtractor, PEImportExtractor, PEResourceExtractor, PEOverlayExtractor, PESectionExtractor, PESignatureExtractor, PEDotNetExtractor, PEInconstistencyTestsExtractor, PEExtraFindings
- **ELF**: ELFFeaturesExtractor, ELFSegmentExtractor, ELFSectionExtractor, ELFDependencyExtractor, ELFSymbolExtractor, ELFImportExtractor, ELFExportExtractor, ELFRelocationExtractor, ELFNotesExtractor
- **Mach-O**: MachOFeaturesExtractor, MachOSegmentExtractor, MachOImportExtractor, MachOExportExtractor, MachODylibExtractor, MachOSignatureExtractor
- **APK**: APKFeaturesExtractor, APKManifestExtractor, APKPermissionsExtractor, APKSignatureExtractor, APKDexExtractor, APKResourceExtractor, APKNativeLibExtractor, APKInconsistencyTestsExtractor
- **Decompilation**: DecompileBinja, DecompileAPK

### Database Schema

The project uses a comprehensive ClickHouse schema defined in `redb/redb_schema.yml` with tables for:
- Basic properties (`redb_basic_properties`)
- PE features (`redb_pe_features`, `redb_pe_imports`, `redb_pe_sections`, etc.)
- Decompiled code (`code_binja_decompiled_functions_content`, `code_binja_decompiled_functions_references`)
- CAPA analysis (`redb_capa`, `redb_capa_capabilities`)

Full schema documentation is available in `docs/database_schema.md`.

### Processing Modes

1. **Analysis Mode**: Extracts features using selected modules
2. **Decompile Mode**: Uses Binary Ninja for code decompilation
3. **S3 Mode**: Fetches samples from S3 storage based on catalog queries
4. **Nomad Job Mode**: Processes single jobs using environment variables for containerized deployment

### Configuration

Environment variables are used for configuration:
- Database connection: `CLICKHOUSE_HOST`, `CLICKHOUSE_PORT`, `CLICKHOUSE_USER`, `CLICKHOUSE_PASSWORD`
- S3 storage: `S3_ENDPOINT`, `S3_ACCESS_KEY`, `S3_SECRET_KEY`
- Processing: `BATCH_SIZE`, `REDB_TIMEOUT`, `DECOMPILE_WORKER_TIMEOUT`
- Nomad jobs: `JOB_ID`, `S3_KEY`, `S3_BUCKET`, `WORKER_TYPE`, `CALLBACK_URL`, `ANALYSIS_MODULES`

## Important Implementation Details

### Multiprocessing
- Uses `spawn` method for multiprocessing to avoid memory issues
- Worker processes have timeout handlers to prevent hanging
- Supports both batch processing and streaming processing modes

### Memory Management
- Implements aggressive garbage collection between batches
- Monitors swap usage and restarts worker pools when needed
- Kills stuck processes automatically

### Error Handling
- Comprehensive logging with per-file context
- Graceful handling of corrupted or unsupported files
- Automatic retry logic for database operations

### Security Context
This is a defensive security tool for malware analysis. It processes potentially malicious files in a controlled environment to extract features for detection and analysis purposes.

## Development Notes

- The codebase is optimized for processing large batches of malware samples
- Extractors are designed to be modular and can be run individually or in combination
- Database schema supports both normalized and denormalized views for different query patterns
- S3 integration allows for scalable processing of large malware repositories