Maite Arratibel

26 papers A 5C 1Misc 1Journal 11Unranked 8
YearRankTypeTitle / Venue / Authors
2025 J jnl
Empir. Softw. Eng.
Pablo Valle, Vincenzo Riccio, Aitor Arrieta, Paolo Tonella, Maite Arratibel
2024 J jnl
Softw. Qual. J.
Iñigo Aldalur, Aitor Arrieta, Aitor Agirre, Goiuria Sagardui, Maite Arratibel
2024 conf
SIGSOFT FSE Companion
Xinyi Wang, Shaukat Ali, Aitor Arrieta, Paolo Arcaini, Maite Arratibel
2024 J jnl
CoRR
Xinyi Wang, Shaukat Ali, Aitor Arrieta, Paolo Arcaini, Maite Arratibel
2024 J jnl
CoRR
Asmar Muqeet, Hassan Sartaj, Aitor Arrieta, Shaukat Ali, Paolo Arcaini, Maite Arratibel, Julie Marie Gjøby, Narasimha Raghavan Veeraragavan, Jan F. Nygård
2024 J jnl
IEEE Trans. Software Eng.
Qinghua Xu, Tao Yue, Shaukat Ali, Maite Arratibel
2023 A conf
ISSTA
Pablo Valle, Aitor Arrieta, Maite Arratibel
2023 J jnl
CoRR
Pablo Valle, Aitor Arrieta, Maite Arratibel
2023 conf
ICSE-SEIP
Pablo Valle, Aitor Arrieta, Maite Arratibel
2023 J jnl
CoRR
Pablo Valle, Aitor Arrieta, Maite Arratibel
2023 J jnl
IEEE Trans. Reliab.
Jon Ayerdi, Pablo Valle, Sergio Segura, Aitor Arrieta, Goiuria Sagardui, Maite Arratibel
2023 J jnl
CoRR
Qinghua Xu, Tao Yue, Shaukat Ali, Maite Arratibel
2023 J jnl
ACM Trans. Softw. Eng. Methodol.
Liping Han, Shaukat Ali, Tao Yue, Aitor Arrieta, Maite Arratibel
2022 conf
ESEC/SIGSOFT FSE
Liping Han, Tao Yue, Shaukat Ali, Aitor Arrieta, Maite Arratibel
2022 A conf
SANER
Aitor Arrieta, Maialen Otaegi, Liping Han, Goiuria Sagardui, Shaukat Ali, Maite Arratibel
2022 conf
GECCO Companion
Jon Ayerdi, Valerio Terragni, Aitor Arrieta, Paolo Tonella, Goiuria Sagardui, Maite Arratibel
2022 J jnl
J. Softw. Evol. Process.
Aitor Gartziandia, Aitor Arrieta, Jon Ayerdi, Miren Illarramendi, Aitor Agirre, Goiuria Sagardui, Maite Arratibel
2022 A conf
ISSRE
Jon Ayerdi, Aitor Arrieta, Ernest Bota Pobee, Maite Arratibel
2022 conf
ESEC/SIGSOFT FSE
Qinghua Xu, Shaukat Ali, Tao Yue, Maite Arratibel
2021 conf
ESEC/SIGSOFT FSE
Jon Ayerdi, Valerio Terragni, Aitor Arrieta, Paolo Tonella, Goiuria Sagardui, Maite Arratibel
2021 conf
ISSRE Workshops
Joritz Galarraga, Aitor Arrieta Marcos, Shaukat Ali, Goiuria Sagardui, Maite Arratibel
2021 conf
ICSA Companion
Aitor Gartziandia, Jon Ayerdi, Aitor Arrieta, Shaukat Ali, Tao Yue, Aitor Agirre, Goiuria Sagardui, Maite Arratibel
2021 C conf
AST
Aitor Arrieta, Jon Ayerdi, Miren Illarramendi, Aitor Agirre, Goiuria Sagardui, Maite Arratibel
2021 Misc conf
SAC
Aitor Gartziandia, Aitor Arrieta, Aitor Agirre, Goiuria Sagardui, Maite Arratibel
2020 A conf
ISSRE
Jon Ayerdi, Sergio Segura, Aitor Arrieta, Goiuria Sagardui, Maite Arratibel
2020 A conf
RE
Jon Ayerdi, Aitor Gartziandia, Aitor Arrieta, Wasif Afzal, Eduard Enoiu, Aitor Agirre, Goiuria Sagardui, Maite Arratibel, Ola Sellin
yara/README.md
← Index yara/README.md markdown
# YARA Rules Directory

This folder contains YARA rules for scanning binary samples.

## Setting Up YARA-Forge Rules

To use the YARA-Forge rules from [https://github.com/YARAHQ/yara-forge](https://github.com/YARAHQ/yara-forge):

```bash
# Download the latest release
cd /path/to/redb/yara
# wget https://github.com/YARAHQ/yara-forge/releases/latest/download/yara-forge-rules-core.zip
wget https://github.com/YARAHQ/yara-forge/releases/latest/download/yara-forge-rules-extended.zip

# Extract rules
# unzip yara-forge-rules-core.zip
unzip yara-forge-rules-extended.zip
```

Available packages:
- `yara-forge-rules-core.zip` - Core rules (~5,000 rules)
- `yara-forge-rules-extended.zip` - Extended rules (~10,000 rules)
- `yara-forge-rules-full.zip` - Full rules (~11,000+ rules)

## Pre-compiling Rules (Recommended for Production)

For large rulesets like YARA-Forge, pre-compiling rules significantly improves startup time:

```bash
# Pre-compile all rules into a single .yarac file
python -m redb.extractors.yara --compile

# Or specify custom paths
python -m redb.extractors.yara --compile --rules-path /path/to/rules --output /path/to/output.yarac
```

This creates `yara/compiled_rules.yarac` which is loaded automatically on subsequent runs.

### Performance Comparison

| Method | First Scan Startup | Subsequent Scans |
|--------|-------------------|------------------|
| Source files (.yar) | ~10-30 seconds (11k rules) | Instant (cached) |
| Pre-compiled (.yarac) | ~1-2 seconds | Instant (cached) |

## Directory Structure

```
yara/
├── README.md
├── .gitkeep
├── compiled_rules.yarac    # (optional) Pre-compiled rules
├── packages/               # YARA-Forge packages
│   └── core/
│       └── *.yar
└── custom/                 # Your custom rules
    └── my_rules.yar
```

Rules are loaded in this priority:
1. `compiled_rules.yarac` (if exists) - fastest
2. All `.yar` and `.yara` files recursively - compiles on first run

## Usage

### Scan with YARA only

```bash
# Scan local files
python start.py --path /path/to/samples -y --repo my_repo --index_prefix redb

# Scan S3 samples
python start.py --s3 --repo bazaar -y --index_prefix redb

# Dry-run (print results instead of storing in ClickHouse)
python start.py --path /path/to/samples -y --dry-run --repo test --index_prefix redb
```

### Scan already-analyzed samples

Run YARA on samples that were previously analyzed (already in `basic_properties`).
Deduplication is handled by the `yara_matches` table — samples already scanned are
automatically excluded before processing begins:

```bash
# Scan all analyzed macho samples with YARA
python start.py --analyzed --magika macho -y --index_prefix redb

# Scan all analyzed PE samples with YARA
python start.py --analyzed --magika pe -y --index_prefix redb

# Scan all analyzed samples (no filetype filter)
python start.py --analyzed -y --index_prefix redb
```

### Partition large YARA runs by date

Combine `--analyzed` with `--range` to partition millions of samples into
manageable batches. Only samples in `basic_properties` AND within the date
range (by `first_seen` in `catalog_samples`) are processed:

```bash
# Scan analyzed PE samples from Feb 2025
python start.py --range 2025-02-01 2025-02-28 --analyzed --magika pebin -y --index_prefix redb

# Scan analyzed PE samples from first week of March 2025
python start.py --range 2025-03-01 2025-03-08 --analyzed --magika pebin -y --index_prefix redb
```

YARA dedup still applies — re-running a range safely skips already-scanned samples.

### Combined Features + YARA

Run feature extraction and YARA scanning together on the same samples:

```bash
# Local files with features + YARA
python start.py --path /path/to/samples --with-yara --repo my_repo --index_prefix redb

# S3 samples with features + YARA
python start.py --s3 --repo bazaar --with-yara --index_prefix redb
```

### Pre-compile Rules

```bash
# Compile and save to default location (yara/compiled_rules.yarac)
python -m redb.extractors.yara --compile

# Compile with custom paths
python -m redb.extractors.yara --compile --rules-path ./my_rules --output ./compiled.yarac
```

### Sync Rules to Database

Before batch scanning, sync rules to ensure all rule metadata is stored:

```bash
# Sync rules to database
python -m redb.extractors.yara --sync-rules

# Sync with custom source collection name
python -m redb.extractors.yara --sync-rules --source-collection yara-forge-core

# Compile and sync in one command
python -m redb.extractors.yara --compile --sync-rules
```

## ClickHouse Table Schema

YARA data uses a **normalized schema** with two tables for efficient storage.

### Matches Table: `yara_matches`

Stores one row per sample-rule match (optimized with binary sha256 and rule_id):

| Column | Type | Description |
|--------|------|-------------|
| sha256 | FixedString(32) | Binary SHA256 (32 bytes, use `hex(sha256)` to display) |
| rule_id | UInt64 | Unique rule identifier (xxHash64 of canonical rule content) |
| rule_name | LowCardinality(String) | YARA rule name (denormalized for convenience) |
| scan_date | DateTime64(3, 'UTC') | Scan timestamp |
| match_strings | Array(String) | Matched string identifiers |

### Rules Table: `yara_rules`

Stores rule metadata once per unique rule (deduplicated by rule_id):

| Column | Type | Description |
|--------|------|-------------|
| rule_id | UInt64 | Unique rule identifier (xxHash64 of canonical rule content) |
| rule_name | String | YARA rule name |
| source_collection | LowCardinality(String) | Source collection (e.g., 'yara-forge-core', 'malpedia') |
| ingested_at | DateTime64(3, 'UTC') | When this rule was ingested |
| rule_text | String | Full rule source code |
| rule_meta | JSON | Rule metadata (author, description, reference, etc.) |
| rule_tags | Array(LowCardinality(String)) | Rule tags |

### Schema Benefits

- **Binary SHA256**: 32 bytes vs 64 bytes (50% storage savings on hash columns)
- **UInt64 rule_id**: Fast joins and lookups via integer key
- **Content-based rule_id**: xxHash64 of canonical rule content (excluding metadata) for deduplication
- **Denormalized rule_name**: Allows queries without joins for common use cases

### Example Queries

```sql
-- Get matches with hex sha256
SELECT
    hex(m.sha256) as sha256,
    m.rule_name,
    m.match_strings
FROM yara_matches m
WHERE m.sha256 = unhex('abc123...')

-- Join with rules for full metadata
SELECT
    hex(m.sha256) as sha256,
    m.rule_name,
    m.match_strings,
    r.rule_meta,
    r.source_collection
FROM yara_matches m
JOIN yara_rules r ON m.rule_id = r.rule_id
WHERE m.sha256 = unhex('abc123...')

-- Find all samples matching a specific rule
SELECT hex(sha256), scan_date
FROM yara_matches
WHERE rule_name = 'APT_Lazarus_Loader'
ORDER BY scan_date DESC
```

## Environment Variables

| Variable | Description | Default |
|----------|-------------|---------|
| `YARA_RULES_PATH` | Override the YARA rules directory | `yara/` |
| `YARA_COMPILED_RULES` | Compiled rules filename | `compiled_rules.yarac` |
| `YARA_SOURCE_COLLECTION` | Default source collection name | `default` |