Karin Quaas

51 papers A* 2A 1B 8C 6Journal 29Unranked 4
YearRankTypeTitle / Venue / Authors
2025 J jnl
Log. Methods Comput. Sci.
Stéphane Demri, Karin Quaas
2023 J jnl
CoRR
Stéphane Demri, Karin Quaas
2023 B conf
CONCUR
Stéphane Demri, Karin Quaas
2023 B conf
JELIA
Stéphane Demri, Karin Quaas
2022 B conf
MFCS
Dominik Peteler, Karin Quaas
2021 J jnl
ACM SIGLOG News
Stéphane Demri, Karin Quaas
2021 A* conf
ICALP
Wojciech Czerwinski, Antoine Mottet, Karin Quaas
2021 J jnl
CoRR
Wojciech Czerwinski, Antoine Mottet, Karin Quaas
2021 J jnl
Theory Comput. Syst.
Antoine Mottet, Karin Quaas
2021 J jnl
Dagstuhl Reports
Thomas Colcombet, Karin Quaas, Michal Skrzypczak
2020 J jnl
Theor. Comput. Sci.
Uli Fahrenberg, Axel Legay, Karin Quaas
2020 J jnl
Inf. Process. Lett.
Martin Fränzle, Karin Quaas, Mahsa Shirmohammadi, James Worrell
2020 J jnl
ACM Trans. Comput. Log.
Shiguang Feng, Claudia Carapelle, Oliver Fernández Gil, Karin Quaas
2019 C conf
ICTAC
Uli Fahrenberg, Axel Legay, Karin Quaas
2019 J jnl
CoRR
Uli Fahrenberg, Axel Legay, Karin Quaas
2019 J jnl
CoRR
Martin Fränzle, Karin Quaas, Mahsa Shirmohammadi, James Worrell
2019 J jnl
CoRR
Antoine Mottet, Karin Quaas
2019 J jnl
ACM Trans. Comput. Log.
Karin Quaas, Mahsa Shirmohammadi
2019 J jnl
Log. Methods Comput. Sci.
Benedikt Bollig, Karin Quaas, Arnaud Sangnier
2019 A conf
STACS
Antoine Mottet, Karin Quaas
2018 J jnl
CoRR
Antoine Mottet, Karin Quaas
2017 J jnl
Acta Cybern.
Zoltán Ésik, Uli Fahrenberg, Axel Legay, Karin Quaas
2017 J jnl
Acta Cybern.
Zoltán Ésik, Uli Fahrenberg, Axel Legay, Karin Quaas
2017 J jnl
Log. Methods Comput. Sci.
Shiguang Feng, Markus Lohrey, Karin Quaas
2017 J jnl
CoRR
Karin Quaas, Mahsa Shirmohammadi, James Worrell
2017 A* conf
LICS
Karin Quaas, Mahsa Shirmohammadi, James Worrell
2017 J jnl
CoRR
Karin Quaas, Mahsa Shirmohammadi
2017 B conf
CONCUR
Benedikt Bollig, Karin Quaas, Arnaud Sangnier
2016 J jnl
CoRR
Benedikt Bollig, Karin Quaas, Arnaud Sangnier
2016 B conf
MFCS
Parvaneh Babari, Karin Quaas, Mahsa Shirmohammadi
2015 C conf
DLT
Shiguang Feng, Markus Lohrey, Karin Quaas
2015 J jnl
Log. Methods Comput. Sci.
Karin Quaas
2014 conf
AFL
Claudia Carapelle, Shiguang Feng, Oliver Fernandez Gil, Karin Quaas
2014 conf
SynCoP
Karin Quaas
2014 J jnl
Theor. Comput. Sci.
Ingmar Meinecke, Karin Quaas
2014 J jnl
CoRR
Shiguang Feng, Markus Lohrey, Karin Quaas
2014 C conf
LATA
Claudia Carapelle, Shiguang Feng, Oliver Fernandez Gil, Karin Quaas
2014 B conf
CONCUR
Karin Quaas
2013 B conf
ATVA
Zoltán Ésik, Uli Fahrenberg, Axel Legay, Karin Quaas
2013 J jnl
CoRR
Zoltán Ésik, Uli Fahrenberg, Axel Legay, Karin Quaas
2013 C conf
LATA
Karin Quaas
2011 J jnl
Theor. Comput. Sci.
Manfred Droste, Karin Quaas
2011 J jnl
Formal Methods Syst. Des.
Karin Quaas
2011 C conf
LATA
Karin Quaas
2011 J jnl
Inf. Process. Lett.
Daniel Kirsten, Karin Quaas
2010
Karin Quaas
2009 conf
FORMATS
Karin Quaas
2009 C conf
Developments in Language Theory
Karin Quaas
2008 B conf
FoSSaCS
Manfred Droste, Karin Quaas
2008 J jnl
Fundam. Informaticae
Parosh Aziz Abdulla, Johann Deneux, Joël Ouaknine, Karin Quaas, James Worrell
2007 conf
FSEN
Parosh Aziz Abdulla, Joël Ouaknine, Karin Quaas, James Worrell
CLAUDE.md
← Index CLAUDE.md markdown
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## Project Overview

REDB (RationalEdge Samples DB) is a malware analysis framework that extracts features from PE (Portable Executable) files and stores them in ClickHouse database for analysis. It provides a comprehensive set of extractors for analyzing binary samples including PE headers, imports, resources, signatures, and decompiled code.

## Common Commands

### Development Setup
```bash
source venv/bin/activate

# Install dependencies
pip install -r requirements.txt

# Run the main application
python start.py --path /path/to/samples --repo sample_repo --index_prefix redb
```

### Analysis Commands
```bash
# Process a single file
python start.py --path /path/to/binary --repo test --index_prefix redb

# Process from S3 storage
python start.py --s3 --repo malpedia --index_prefix redb

# Process from S3 storage but only a subset of a specific repository
python start.py --s3 --repo "vx-itw" --s3-notes "ITW.0138" --index_prefix redb

# Run only decompilation
python start.py --path /path/to/binary --repo test --index_prefix redb --decompile

# Run specific modules
python start.py --path /path/to/binary --repo test --index_prefix redb --modules "BasicPropertiesExtractor,PEFeaturesExtractor"

# Run as Nomad job (for containerized deployment)
python start.py --nomad-job
```

### Testing
There are no formal unit tests. Testing is done by running the extractors on sample files in the `test_files/` directory.

## Architecture Overview

### Core Components

1. **Ingestor (`redb/ingestor.py`)**: Main orchestrator that handles file processing, multiprocessing, and coordinates extractors
2. **Extractors (`redb/extractors/`)**: Modular analysis components that extract specific features
3. **Database Exporters (`redb/extractors/database_exporters.py`)**: Handle data export to ClickHouse
4. **Settings (`redb/settings/`)**: Configuration management for database connections

### Extractor Architecture

All extractors inherit from the base `Extractor` class and implement:
- `extract()`: Main analysis logic
- `prepare_export_data()`: Format data for database export
- `get_clickhouse_table()`: Return target table name

Available extractors:
- **General**: BasicPropertiesExtractor, HashExtractor, DIEExtractor, CAPAExtractor
- **PE-specific**: PEFeaturesExtractor, PEImportExtractor, PEResourceExtractor, PEOverlayExtractor, PESectionExtractor, PESignatureExtractor, PEDotNetExtractor, PEInconstistencyTestsExtractor, PEExtraFindings
- **ELF**: ELFFeaturesExtractor, ELFSegmentExtractor, ELFSectionExtractor, ELFDependencyExtractor, ELFSymbolExtractor, ELFImportExtractor, ELFExportExtractor, ELFRelocationExtractor, ELFNotesExtractor
- **Mach-O**: MachOFeaturesExtractor, MachOSegmentExtractor, MachOImportExtractor, MachOExportExtractor, MachODylibExtractor, MachOSignatureExtractor
- **APK**: APKFeaturesExtractor, APKManifestExtractor, APKPermissionsExtractor, APKSignatureExtractor, APKDexExtractor, APKResourceExtractor, APKNativeLibExtractor, APKInconsistencyTestsExtractor
- **Decompilation**: DecompileBinja, DecompileAPK

### Database Schema

The project uses a comprehensive ClickHouse schema defined in `redb/redb_schema.yml` with tables for:
- Basic properties (`redb_basic_properties`)
- PE features (`redb_pe_features`, `redb_pe_imports`, `redb_pe_sections`, etc.)
- Decompiled code (`code_binja_decompiled_functions_content`, `code_binja_decompiled_functions_references`)
- CAPA analysis (`redb_capa`, `redb_capa_capabilities`)

Full schema documentation is available in `docs/database_schema.md`.

### Processing Modes

1. **Analysis Mode**: Extracts features using selected modules
2. **Decompile Mode**: Uses Binary Ninja for code decompilation
3. **S3 Mode**: Fetches samples from S3 storage based on catalog queries
4. **Nomad Job Mode**: Processes single jobs using environment variables for containerized deployment

### Configuration

Environment variables are used for configuration:
- Database connection: `CLICKHOUSE_HOST`, `CLICKHOUSE_PORT`, `CLICKHOUSE_USER`, `CLICKHOUSE_PASSWORD`
- S3 storage: `S3_ENDPOINT`, `S3_ACCESS_KEY`, `S3_SECRET_KEY`
- Processing: `BATCH_SIZE`, `REDB_TIMEOUT`, `DECOMPILE_WORKER_TIMEOUT`
- Nomad jobs: `JOB_ID`, `S3_KEY`, `S3_BUCKET`, `WORKER_TYPE`, `CALLBACK_URL`, `ANALYSIS_MODULES`

## Important Implementation Details

### Multiprocessing
- Uses `spawn` method for multiprocessing to avoid memory issues
- Worker processes have timeout handlers to prevent hanging
- Supports both batch processing and streaming processing modes

### Memory Management
- Implements aggressive garbage collection between batches
- Monitors swap usage and restarts worker pools when needed
- Kills stuck processes automatically

### Error Handling
- Comprehensive logging with per-file context
- Graceful handling of corrupted or unsupported files
- Automatic retry logic for database operations

### Security Context
This is a defensive security tool for malware analysis. It processes potentially malicious files in a controlled environment to extract features for detection and analysis purposes.

## Development Notes

- The codebase is optimized for processing large batches of malware samples
- Extractors are designed to be modular and can be run individually or in combination
- Database schema supports both normalized and denormalized views for different query patterns
- S3 integration allows for scalable processing of large malware repositories