Olli Yli-Harja

89 papers A 1B 2C 1Journal 66Unranked 19
YearRankTypeTitle / Venue / Authors
2025 J jnl
npj Digit. Medicine
Yue Zhang, Jonathan Martinez, Bahar Tercan, Heikki Kuusanmäki, Frank Emmert-Streib, Jerome G. Chandraseelan, Amer Farea, Olli Yli-Harja, Caroline A. Heckman, David L. Gibbs, Vesteinn Thorsson, Ilya Shmulevich, Mika Kontro, Guangrong Qin, Boris Aguilar
2025 J jnl
npj Digit. Medicine
Frank Emmert-Streib, Seppo Parkkila, Reinhard Laubenbacher, Arto Mannermaa, Leroy Hood, Olli Yli-Harja
2024 J jnl
IEEE Access
Frank Emmert-Streib, Shailesh Tripathi, Amer Farea, Olli Yli-Harja
2022 J jnl
CoRR
Syeda Sakira Hassan, Rahul Mangayil, Tommi Aho, Olli Yli-Harja, Matti Karp
2020 J jnl
CoRR
Frank Emmert-Streib, Olli Yli-Harja, Matthias Dehmer
2020 J jnl
Frontiers Artif. Intell.
Frank Emmert-Streib, Olli Yli-Harja, Matthias Dehmer
2020 J jnl
CoRR
Frank Emmert-Streib, Olli Yli-Harja, Matthias Dehmer
2020 J jnl
WIREs Data Mining Knowl. Discov.
Frank Emmert-Streib, Olli Yli-Harja, Matthias Dehmer
2019 J jnl
Mach. Learn. Knowl. Extr.
Aliyu Musa, Matthias Dehmer, Olli Yli-Harja, Frank Emmert-Streib
2018 J jnl
Briefings Bioinform.
Aliyu Musa, Laleh Soltan Ghoraie, Shu-Dong Zhang, Galina V. Glazko, Olli Yli-Harja, Matthias Dehmer, Benjamin Haibe-Kains, Frank Emmert-Streib
2018 J jnl
Frontiers Appl. Math. Stat.
Frank Emmert-Streib, Shailesh Tripathi, Olli Yli-Harja, Matthias Dehmer
2017 J jnl
Briefings Bioinform.
Aliyu Musa, Laleh Soltan Ghoraie, Shu-Dong Zhang, Galina V. Glazko, Olli Yli-Harja, Matthias Dehmer, Benjamin Haibe-Kains, Frank Emmert-Streib
2017 J jnl
BMC Bioinform.
Shailesh Tripathi, Jason Lloyd-Price, Andre S. Ribeiro, Olli Yli-Harja, Matthias Dehmer, Frank Emmert-Streib
2016 conf
WIVACE
Olli Yli-Harja, Frank Emmert-Streib, Jari Yli-Hietanen
2016 J jnl
Nano Commun. Networks
Olga Anufrieva, Adrien Sala, Olli Yli-Harja, Meenakshisundaram Kandhavelu
2015 conf
EUSIPCO
Septimia Sârbu, Ilya Shmulevich, Olli Yli-Harja, Matti Nykter, Juha Kesseli
2013 J jnl
Pattern Recognit.
Muhammad Farhan, Olli Yli-Harja, Antti Niemistö
2013 J jnl
BMC Bioinform.
Sharif Chowdhury, Meenakshisundaram Kandhavelu, Olli Yli-Harja, Andre S. Ribeiro
2013 J jnl
BMC Syst. Biol.
Antti Ylipää, Olli Yli-Harja, Wei Zhang, Matti Nykter
2013 J jnl
BMC Syst. Biol.
Antti Häkkinen, Huy Tran, Olli Yli-Harja, Brian P. Ingalls, Andre S. Ribeiro
2013 conf
Image Processing: Algorithms and Systems
Muhammad Farhan, Pekka Ruusuvuori, Mario Emmenlauer, Pauli Rämö, Olli Yli-Harja, Christoph Dehio
2013 conf
MLSP
Muhammad Farhan, Antti Larjo, Olli Yli-Harja, Tommi Aho
2013 J jnl
BMC Bioinform.
Muhammad Farhan, Pekka Ruusuvuori, Mario Emmenlauer, Pauli Rämö, Christoph Dehio, Olli Yli-Harja
2013 J jnl
BMC Bioinform.
Teppo Annila, Eero Lihavainen, Ines J. Marques, Darren R. Williams, Olli Yli-Harja, Andre S. Ribeiro
2012 J jnl
Biosyst.
Meenakshisundaram Kandhavelu, Eero Lihavainen, Anantha Barathi Muthukrishnan, Olli Yli-Harja, Andre Sanches Ribeiro
2012 conf
CMSB
Ilya Potapov, Jarno Mäkelä, Olli Yli-Harja, Andre S. Ribeiro
2011 J jnl
BMC Bioinform.
Kirsti Laurila, Bodil Oster, Claus L. Andersen, Philippe Lamy, Torben F. Ørntoft, Olli Yli-Harja, Carsten Wiuf
2011 J jnl
Bioinform.
Jarno Mäkelä, Heikki Huttunen, Meenakshisundaram Kandhavelu, Olli Yli-Harja, Andre S. Ribeiro
2011 J jnl
Biosyst.
Andre S. Ribeiro, Jason Lloyd-Price, Sharif Chowdhury, Olli Yli-Harja
2011 J jnl
Trans. Comp. Sys. Biology
Antti Häkkinen, Fred G. Biddle, Olli-Pekka Smolander, Olli Yli-Harja, Andre S. Ribeiro
2011 J jnl
BMC Syst. Biol.
Meenakshisundaram Kandhavelu, Henrik Mannerström, Abhishekh Gupta, Antti Häkkinen, Jason Lloyd-Price, Olli Yli-Harja, Andre S. Ribeiro
2011 J jnl
EURASIP J. Bioinform. Syst. Biol.
Henrik Mannerström, Olli Yli-Harja, Andre S. Ribeiro
2011 J jnl
Frontiers Comput. Neurosci.
Tuomo Mäki-Marttunen, Jugoslava Acimovic, Matti Nykter, Juha Kesseli, Keijo Ruohonen, Olli Yli-Harja, Marja-Leena Linne
2011 J jnl
BMC Bioinform.
Jarno Mäkelä, Jason Lloyd-Price, Olli Yli-Harja, Andre S. Ribeiro
2010 J jnl
Comput. Biol. Chem.
Andre S. Ribeiro, Antti Häkkinen, Shannon Healy, Olli Yli-Harja
2010 J jnl
PLoS Comput. Biol.
Tiina Rajala, Antti Häkkinen, Shannon Healy, Olli Yli-Harja, Andre S. Ribeiro
2010 J jnl
BMC Bioinform.
Pekka Ruusuvuori, Tarmo Äijö, Sharif Chowdhury, Cecilia Garmendia-Torres, Jyrki Selinummi, Mirko Birbaumer, Aimée M. Dudley, Lucas Pelkmans, Olli Yli-Harja
2010 B conf
ICIP
Hannu Oinonen, Heikki Forsvik, Pekka Ruusuvuori, Olli Yli-Harja, Ville Voipio, Heikki Huttunen
2010 J jnl
BMC Syst. Biol.
Sharif Chowdhury, Jason Lloyd-Price, Olli-Pekka Smolander, Wayne C. V. Baici, Timothy R. Hughes, Olli Yli-Harja, Gordon Chua, Andre S. Ribeiro
2010 J jnl
EURASIP J. Adv. Signal Process.
Xiaofeng Dai, Olli Yli-Harja, Harri Lähdesmäki
2009 J jnl
BMC Bioinform.
Xiaofeng Dai, Timo Erkkilä, Olli Yli-Harja, Harri Lähdesmäki
2009 J jnl
J. Comput. Biol.
Andre S. Ribeiro, Olli-Pekka Smolander, Tiina Rajala, Antti Häkkinen, Olli Yli-Harja
2009 J jnl
Bioinform.
Xiaofeng Dai, Olli Yli-Harja, Andre S. Ribeiro
2008 conf
EUSIPCO
Pekka Ruusuvuori, Antti Lehmussola, Jyrki Selinummi, Tiina Rajala, Heikki Huttunen, Olli Yli-Harja
2008 B conf
ICPR
Pekka Ruusuvuori, Jenni Seppälä, Timo Erkkilä, Antti Lehmussola, Jaakko A. Puhakka, Olli Yli-Harja
2008 J jnl
PLoS Comput. Biol.
Antti Saarinen, Marja-Leena Linne, Olli Yli-Harja
2008 J jnl
Proc. IEEE
Antti Lehmussola, Pekka Ruusuvuori, Jyrki Selinummi, Tiina Rajala, Olli Yli-Harja
2008 J jnl
Pattern Recognit. Lett.
Alireza Akhbardeh, Nikhil, Perttu E. Koskinen, Olli Yli-Harja
2008 conf
BIOTECHNO
Xiaofeng Dai, Harri Lähdesmäki, Olli Yli-Harja
2007 J jnl
IEEE Trans. Medical Imaging
Antti Lehmussola, Pekka Ruusuvuori, Jyrki Selinummi, Heikki Huttunen, Olli Yli-Harja
2007 J jnl
EURASIP J. Bioinform. Syst. Biol.
Antti Niemistö, Matti Nykter, Tommi Aho, Henna Jalovaara, Kalle Marjanen, Miika Ahdesmäki, Pekka Ruusuvuori, Mikko Tiainen, Marja-Leena Linne, Olli Yli-Harja
2007 J jnl
BMC Syst. Biol.
Tommi Aho, Olli-Pekka Smolander, Jari Niemi, Olli Yli-Harja
2007 J jnl
BMC Bioinform.
Miika Ahdesmäki, Harri Lähdesmäki, Andrew Gracey, Ilya Shmulevich, Olli Yli-Harja
2006 J jnl
Neurocomputing
Antti Pettinen, Olli Yli-Harja, Marja-Leena Linne
2006 J jnl
Bioinform.
Antti Lehmussola, Pekka Ruusuvuori, Olli Yli-Harja
2006 conf
EMBC
Antti Niemistö, Jyrki Selinummi, Ramsey Saleem, Ilya Shmulevich, John D. Aitchison, Olli Yli-Harja
2006 C conf
Image Processing: Algorithms and Systems
Pekka Ruusuvuori, Antti Lehmussola, Olli Yli-Harja
2006 J jnl
Neurocomputing
Antti Saarinen, Marja-Leena Linne, Olli Yli-Harja
2006 J jnl
Signal Process.
Harri Lähdesmäki, Sampsa Hautaniemi, Ilya Shmulevich, Olli Yli-Harja
2006 J jnl
BMC Bioinform.
Matti Nykter, Tommi Aho, Miika Ahdesmäki, Pekka Ruusuvuori, Antti Lehmussola, Olli Yli-Harja
2006 conf
EMBC
Jyrki Selinummi, Pekka Ruusuvuori, Antti Lehmussola, Heikki Huttunen, Olli Yli-Harja, Riitta Miettinen
2005 conf
ICIP (1)
Antti Lehmussola, Pekka Ruusuvuori, Olli Yli-Harja
2005 J jnl
BMC Bioinform.
Miika Ahdesmäki, Harri Lähdesmäki, Ronald K. Pearson, Heikki Huttunen, Olli Yli-Harja
2005 J jnl
IEEE Trans. Medical Imaging
Antti Niemistö, Valerie Dunmire, Olli Yli-Harja, Wei Zhang, Ilya Shmulevich
2005 J jnl
Bioinform.
Antti Pettinen, Tommi Aho, Olli-Pekka Smolander, Tiina Manninen, Antti Saarinen, Kaisa-Leena Taattola, Olli Yli-Harja, Marja-Leena Linne
2004 J jnl
J. Frankl. Inst.
Sampsa Hautaniemi, Markus Ringnér, Päivikki Kauraniemi, Reija Autio, Henrik Edgren, Olli Yli-Harja, Jaakko Astola, Anne Kallioniemi, Olli-Pekka Kallioniemi
2004 J jnl
EURASIP J. Adv. Signal Process.
Antti Niemistö, Ilya Shmulevich, Vladimir Vasilyevich Lukin, Alexander N. Dolia, Olli Yli-Harja
2003 J jnl
Bioinform.
Sampsa Hautaniemi, Henrik Edgren, Petri Vesanen, Maija Wolf, Anna-Kaarina Järvinen, Olli Yli-Harja, Jaakko Astola, Olli-P. Kallioniemi, Outi Monni
2003 J jnl
Mach. Learn.
Sampsa Hautaniemi, Olli Yli-Harja, Jaakko Astola, Päivikki Kauraniemi, Anne Kallioniemi, Maija Wolf, Jimmy Ruiz, Spyro Mousses, Olli-P. Kallioniemi
2003 J jnl
Bioinform.
Reija Autio, Sampsa Hautaniemi, Päivikki Kauraniemi, Olli Yli-Harja, Jaakko Astola, Maija Wolf, Anne Kallioniemi
2003 A conf
SDM
Ronald K. Pearson, Harri Lähdesmäki, Heikki Huttunen, Olli Yli-Harja
2003 J jnl
Signal Process.
Harri Lähdesmäki, Heikki Huttunen, Tommi Aho, Marja-Leena Linne, Jari Niemi, Juha Kesseli, Ronald K. Pearson, Olli Yli-Harja
2003 conf
Image Processing: Algorithms and Systems
Antti Niemistö, Tommi Aho, Henna Thesleff, Mikko Tiainen, Kalle Marjanen, Marja-Leena Linne, Olli Yli-Harja
2003 conf
Image Processing: Algorithms and Systems
Jari Niemi, Kalle Marjanen, Heimo Ihalainen, Olli Yli-Harja
2003 conf
NNSP
John A. Berger, Sampsa Hautaniemi, Henrik Edgren, Outi Monni, Sanjit K. Mitra, Olli Yli-Harja, Jaakko Astola
2003 J jnl
Mach. Learn.
Harri Lähdesmäki, Ilya Shmulevich, Olli Yli-Harja
2003 conf
Image Processing: Algorithms and Systems
Heikki Huttunen, Pertti T. Koivisto, Antti Niemistö, Olli Yli-Harja
2002 conf
Image Processing: Algorithms and Systems
Petri Vesanen, Mikko Tiainen, Olli Yli-Harja
2002 conf
Image Processing: Algorithms and Systems
Kalle Marjanen, Heimo Ihalainen, Heikki Huttunen, Olli Yli-Harja
2002 J jnl
IEEE Trans. Signal Process.
Ilya Shmulevich, Olli Yli-Harja, Jaakko Astola, Aleksey D. Korshunov
2001 J jnl
Signal Process.
Pertti Koivisto, Olli Yli-Harja, Antti Niemistö, Ilya Shmulevich
2001 J jnl
Comput. Humanit.
Ilya Shmulevich, Olli Yli-Harja, Edward J. Coyle, Dirk-Jan Povel, Kjell Lemström
2001 J jnl
Signal Process.
Olli Yli-Harja, Pertti Koivisto, J. Andrew Bangham, Gavin C. Cawley, Richard W. Harvey, Ilya Shmulevich
2000 conf
EUSIPCO
Antti Niemistö, Olli Yli-Harja, Antti Valmari, Pertti Koivisto, Ilya Shmulevich
1999 conf
NSIP
Olli Yli-Harja, J. Andrew Bangham, Richard W. Harvey, Richard V. Aldridge
1999 conf
NSIP
Olli Yli-Harja, Ilya Shmulevich, Kjell Lemström
1999 J jnl
IEEE Signal Process. Lett.
Ilya Shmulevich, Olli Yli-Harja, Karen O. Egiazarian, Jaakko Astola
1994 J jnl
IEEE Signal Process. Lett.
Olli Yli-Harja
1991 J jnl
IEEE Trans. Signal Process.
Olli Yli-Harja, Jaakko Astola, Yrjö Neuvo
redb/extractors/yara.py
← Index redb/extractors/yara.py python
"""
YARA Extractor - Scans binary files with YARA rules and stores matches in ClickHouse.

This extractor uses the yara-x library to scan samples against a collection of
YARA rules located in the 'yara/' folder at the project root.

Supports both:
- Pre-compiled rules (.yarac) for faster loading
- Source rules (.yar/.yara) compiled on-the-fly

Schema Design:
- yara_rules: Rule metadata stored once per unique rule (deduplicated by rule_id)
- yara_matches: Sample-rule matches with binary sha256 for efficiency
"""
import inspect
import json
import os
import re
import threading
import time
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple

import xxhash
import yara_x

from redb.extractors.enum import Tag
from redb.extractors.extractor import Extractor
from redb.models.dataclasses import YaraMatch, YaraRule


# Default compiled rules filename (can be overridden via YARA_COMPILED_RULES env var)
COMPILED_RULES_FILENAME = os.getenv("YARA_COMPILED_RULES", "compiled_rules.yarac")

# Default source collection name
DEFAULT_SOURCE_COLLECTION = os.getenv("YARA_SOURCE_COLLECTION", "default")


def canonicalize_rule(rule_text: str) -> str:
    """
    Canonicalize YARA rule text for consistent hashing.

    Strips comments and normalizes whitespace, but excludes metadata
    so rules with same logic but different metadata get the same ID.
    """
    # Strip single-line comments
    text = re.sub(r'//.*$', '', rule_text, flags=re.MULTILINE)
    # Strip multi-line comments
    text = re.sub(r'/\*.*?\*/', '', text, flags=re.DOTALL)
    # Strip meta section (keep only strings/condition)
    text = re.sub(r'meta\s*:\s*[^}]+(?=strings|condition|})', '', text, flags=re.DOTALL)
    # Normalize whitespace
    text = ' '.join(text.split())
    return text


def generate_rule_id(rule_text: str) -> int:
    """
    Generate a unique rule_id from canonicalized rule content.

    Returns:
        UInt64 hash of the canonical rule content
    """
    canonical = canonicalize_rule(rule_text)
    return xxhash.xxh64(canonical.encode('utf-8')).intdigest()


def sha256_hex_to_binary(hex_str: str) -> bytes:
    """Convert SHA256 hex string to binary (32 bytes)."""
    return bytes.fromhex(hex_str)


def sha256_binary_to_hex(binary: bytes) -> str:
    """Convert SHA256 binary (32 bytes) to hex string."""
    return binary.hex()


def parse_yara_rules(file_content: str) -> List[Dict[str, Any]]:
    """
    Parse individual YARA rules from file content.

    Handles files with multiple rules and extracts:
    - rule_name: The rule identifier
    - rule_tags: List of tags (from 'rule Name : tag1 tag2 {')
    - rule_meta: Dict of metadata key-value pairs
    - rule_text: Full rule source code

    Args:
        file_content: Raw content of a .yar/.yara file

    Returns:
        List of dicts, each containing rule_name, rule_tags, rule_meta, rule_text
    """
    rules = []

    # Pattern to match rule declarations: rule Name or rule Name : tags
    # We need to find rule boundaries by tracking braces
    rule_pattern = re.compile(
        r'(?:^|\n)\s*((?:private\s+|global\s+)*rule\s+(\w+)\s*(?::\s*([^{]*))?\s*\{)',
        re.MULTILINE
    )

    matches = list(rule_pattern.finditer(file_content))

    for i, match in enumerate(matches):
        rule_start = match.start(1)  # Start of 'rule ...'
        rule_name = match.group(2)
        tags_str = match.group(3) or ""
        rule_tags = [t.strip() for t in tags_str.split() if t.strip()]

        # Find the matching closing brace by counting braces
        brace_count = 0
        rule_end = match.end()
        in_string = False
        escape_next = False

        for j, char in enumerate(file_content[match.end() - 1:], start=match.end() - 1):
            if escape_next:
                escape_next = False
                continue
            if char == '\\':
                escape_next = True
                continue
            if char == '"' and not escape_next:
                in_string = not in_string
                continue
            if in_string:
                continue
            if char == '{':
                brace_count += 1
            elif char == '}':
                brace_count -= 1
                if brace_count == 0:
                    rule_end = j + 1
                    break

        rule_text = file_content[rule_start:rule_end].strip()

        # Extract metadata from rule
        meta = {}
        meta_match = re.search(
            r'meta\s*:\s*(.*?)(?=strings\s*:|condition\s*:|$)',
            rule_text,
            re.DOTALL
        )
        if meta_match:
            meta_text = meta_match.group(1)
            for line in meta_text.strip().split('\n'):
                if '=' in line:
                    key, _, value = line.partition('=')
                    key = key.strip()
                    value = value.strip().strip('"\'')
                    if key and not key.startswith('//'):
                        meta[key] = value

        rules.append({
            'rule_name': rule_name,
            'rule_tags': rule_tags,
            'rule_meta': meta,
            'rule_text': rule_text,
        })

    return rules


class YaraBatchInsertBuffer:
    """
    Thread-safe buffer for batching YARA match inserts.

    Collects matches from multiple samples and flushes to ClickHouse
    when batch_size is reached or flush_interval expires.
    """

    def __init__(
        self,
        exporter,
        index_prefix: str,
        batch_size: int = 1000,
        flush_interval: int = 30,
    ):
        self.exporter = exporter
        self.index_prefix = index_prefix
        self.batch_size = batch_size
        self.flush_interval = flush_interval

        self.matches_buffer: List[List] = []
        self.rules_buffer: Dict[int, List] = {}  # rule_id -> rule_data (deduplicated)
        self.lock = threading.Lock()
        self.last_flush = time.time()

        # Start background flush timer
        self._stop_timer = False
        self._timer_thread = threading.Thread(target=self._flush_timer, daemon=True)
        self._timer_thread.start()

    def add(self, sha256_binary: bytes, matches: List[Tuple], rules: Dict[int, List]) -> None:
        """
        Add matches and rules to buffer.

        Args:
            sha256_binary: Binary SHA256 (32 bytes)
            matches: List of match tuples (sha256_binary, rule_id, rule_name, scan_date, match_strings)
            rules: Dict of rule_id -> rule_data tuples
        """
        with self.lock:
            self.matches_buffer.extend(matches)
            # Merge rules (deduplicated by rule_id)
            for rule_id, rule_data in rules.items():
                if rule_id not in self.rules_buffer:
                    self.rules_buffer[rule_id] = rule_data

            if len(self.matches_buffer) >= self.batch_size:
                self._flush_locked()

    def _flush_locked(self) -> None:
        """Flush buffer (must hold lock)."""
        if not self.matches_buffer:
            return

        try:
            # Insert rules first (deduplicated)
            if self.rules_buffer:
                rules_data = list(self.rules_buffer.values())
                self.exporter.batch_insert(
                    table='yara_rules',
                    rows=rules_data,
                    column_names=[
                        'rule_id', 'rule_name', 'source_collection',
                        'ingested_at', 'rule_text', 'rule_meta', 'rule_tags'
                    ],
                    column_type_names=[
                        'UInt64', 'String', 'LowCardinality(String)',
                        "DateTime64(3, 'UTC')", 'String', 'JSON', 'Array(LowCardinality(String))'
                    ]
                )

            # Insert matches
            self.exporter.batch_insert(
                table='yara_matches',
                rows=self.matches_buffer,
                column_names=[
                    'sha256', 'rule_id', 'rule_name',
                    'scan_date', 'match_strings'
                ],
                column_type_names=[
                    'FixedString(32)', 'UInt64', 'LowCardinality(String)',
                    "DateTime64(3, 'UTC')", 'Array(String)'
                ]
            )

            print(f"[YARA] Flushed {len(self.matches_buffer)} matches and {len(self.rules_buffer)} rules")

        except Exception as e:
            print(f"[YARA] Error flushing batch: {e}")

        # Clear buffers
        self.matches_buffer = []
        self.rules_buffer = {}
        self.last_flush = time.time()

    def flush(self) -> None:
        """Public flush (acquires lock)."""
        with self.lock:
            self._flush_locked()

    def _flush_timer(self) -> None:
        """Background timer for periodic flushes."""
        while not self._stop_timer:
            time.sleep(5)  # Check every 5 seconds
            with self.lock:
                if time.time() - self.last_flush > self.flush_interval and self.matches_buffer:
                    self._flush_locked()

    def stop(self) -> None:
        """Stop the background timer and flush remaining data."""
        self._stop_timer = True
        self.flush()


class YaraExtractor(Extractor):
    """
    Extractor that scans binary files with YARA rules.

    The YARA rules are loaded from the 'yara/' folder at the project root.
    Each matching rule is stored in ClickHouse with its metadata.

    Supports pre-compiled rules for faster loading:
    - If 'yara/compiled_rules.yarac' exists, it will be loaded directly
    - Otherwise, all .yar/.yara files are compiled and cached in memory
    - Use YaraExtractor.compile_and_save() to pre-compile rules
    """

    # Class-level cache for compiled YARA rules
    _compiled_rules = None
    _rules_path = None

    def __init__(
        self,
        filepath: str,
        log: Any,
        exporters: Optional[List] = None,
        index_prefix: Optional[str] = None,
        known_benign: bool = False,
        known_malicious: bool = False,
    ):
        super().__init__(
            filepath,
            log,
            exporters,
            index_prefix,
            known_benign,
            known_malicious,
        )
        self.matches: List[YaraMatch] = []
        self.rules: Dict[str, YaraRule] = {}  # rule_name -> YaraRule (deduplicated)

    @classmethod
    def get_yara_rules_path(cls) -> Path:
        """Get the path to the YARA rules directory."""
        # Default to 'yara/' in project root
        project_root = Path(__file__).parent.parent.parent
        default_path = project_root / "yara"

        # Allow override via environment variable
        rules_path = os.getenv("YARA_RULES_PATH", str(default_path))
        return Path(rules_path)

    @classmethod
    def get_compiled_rules_path(cls) -> Path:
        """Get the path to the pre-compiled rules file."""
        return cls.get_yara_rules_path() / COMPILED_RULES_FILENAME

    @classmethod
    def load_compiled_rules(cls) -> Optional[yara_x.Rules]:
        """
        Load pre-compiled YARA rules from .yarac file.

        Returns:
            Compiled Rules object or None if file doesn't exist
        """
        compiled_path = cls.get_compiled_rules_path()
        if not compiled_path.exists():
            return None

        try:
            with open(compiled_path, "rb") as f:
                rules = yara_x.Rules.deserialize_from(f)
            # Only log in debug mode (non-prod) to avoid spamming in bulk processing
            if os.getenv("SERVER_ENV") != "prod":
                print(f"[DEBUG] Loaded pre-compiled YARA rules from {compiled_path}")
            return rules
        except Exception as e:
            print(f"[WARNING] Failed to load compiled rules from {compiled_path}: {e}")
            return None

    @classmethod
    def compile_rules_from_source(cls) -> Optional[yara_x.Rules]:
        """
        Compile all YARA rules from source .yar/.yara files.

        Returns:
            Compiled Rules object or None if no rules found
        """
        rules_path = cls.get_yara_rules_path()

        if not rules_path.exists():
            return None

        # Find all .yar and .yara files recursively
        rule_files = []
        for ext in ["*.yar", "*.yara"]:
            rule_files.extend(rules_path.rglob(ext))

        if not rule_files:
            return None

        print(f"[INFO] Compiling {len(rule_files)} YARA rule files from {rules_path}")

        # Compile all rules using yara-x compiler
        compiler = yara_x.Compiler()
        compiled_count = 0
        failed_count = 0

        for rule_file in rule_files:
            try:
                with open(rule_file, "r", encoding="utf-8") as f:
                    rule_content = f.read()
                # Use the relative path from rules_path as namespace
                namespace = str(rule_file.relative_to(rules_path).parent)
                if namespace == ".":
                    namespace = "default"
                compiler.new_namespace(namespace)
                compiler.add_source(rule_content)
                compiled_count += 1
            except Exception as e:
                # Log warning but continue with other rules
                print(f"[WARNING] Failed to compile YARA rule {rule_file}: {e}")
                failed_count += 1
                continue

        try:
            rules = compiler.build()
            print(f"[INFO] Successfully compiled {compiled_count} rule files ({failed_count} failed)")
            return rules
        except Exception as e:
            print(f"[ERROR] Failed to build YARA rules: {e}")
            return None

    @classmethod
    def compile_rules(cls, force_reload: bool = False) -> Optional[yara_x.Rules]:
        """
        Get compiled YARA rules, loading from cache, .yarac file, or compiling from source.

        Priority:
        1. Return cached rules if available
        2. Load pre-compiled .yarac file if it exists
        3. Compile from source .yar/.yara files

        Args:
            force_reload: If True, ignore cache and reload rules

        Returns:
            Compiled YARA rules or None if no rules found
        """
        rules_path = cls.get_yara_rules_path()

        # Return cached rules if available and path hasn't changed
        if (
            cls._compiled_rules is not None
            and cls._rules_path == rules_path
            and not force_reload
        ):
            return cls._compiled_rules

        # Try loading pre-compiled rules first
        rules = cls.load_compiled_rules()

        # If no pre-compiled rules, compile from source
        if rules is None:
            rules = cls.compile_rules_from_source()

        # Cache the rules
        if rules is not None:
            cls._compiled_rules = rules
            cls._rules_path = rules_path

        return rules

    @classmethod
    def compile_and_save(cls, output_path: Optional[Path] = None) -> bool:
        """
        Compile all YARA rules from source and save to a .yarac file.

        This is useful for pre-compiling rules for faster loading in production.

        Args:
            output_path: Path to save compiled rules (default: yara/compiled_rules.yarac)

        Returns:
            True if successful, False otherwise
        """
        if output_path is None:
            output_path = cls.get_compiled_rules_path()

        # Force compile from source
        rules = cls.compile_rules_from_source()
        if rules is None:
            print("[ERROR] No rules to compile")
            return False

        try:
            # Ensure parent directory exists
            output_path.parent.mkdir(parents=True, exist_ok=True)

            with open(output_path, "wb") as f:
                rules.serialize_into(f)

            print(f"[INFO] Saved compiled YARA rules to {output_path}")
            return True
        except Exception as e:
            print(f"[ERROR] Failed to save compiled rules: {e}")
            return False

    def _extract_yara_matches(self) -> Tuple[List[YaraMatch], Dict[str, YaraRule]]:
        """
        Scan the binary with compiled YARA rules.

        Returns:
            Tuple of (matches list, rules dict) for each matching rule
        """
        self.log.debug(inspect.currentframe().f_code.co_name)

        compiled_rules = self.compile_rules()
        if compiled_rules is None:
            self.log.warning("No YARA rules found or compiled")
            return [], {}

        matches = []
        rules = {}

        try:
            # Scan the binary data
            scan_results = compiled_rules.scan(self.binary)

            # Process each matching rule
            for rule in scan_results.matching_rules:
                rule_name = rule.identifier

                # Extract rule metadata and store rule (deduplicated by name)
                if rule_name not in rules:
                    meta = {}
                    for identifier, value in rule.metadata:
                        meta[identifier] = str(value)
                    rules[rule_name] = YaraRule(
                        rule_name=rule_name,
                        rule_tags=list(rule.tags),
                        rule_meta=meta,
                    )

                # Extract matched string identifiers
                matched_strings = []
                for pattern in rule.patterns:
                    for match in pattern.matches:
                        matched_strings.append(pattern.identifier)

                # Remove duplicates from matched strings
                matched_strings = list(set(matched_strings))

                yara_match = YaraMatch(
                    rule_name=rule_name,
                    match_strings=matched_strings,
                )
                matches.append(yara_match)

                self.log.debug(
                    f"YARA match: {rule_name} (tags: {list(rule.tags)})"
                )

        except Exception as e:
            self.log.error(f"Error scanning with YARA: {e}")
            return [], {}

        return matches, rules

    def extract(self) -> Optional[List[YaraMatch]]:
        """
        Extract YARA matches from the binary.

        Returns:
            List of YaraMatch objects or None on error
        """
        self.log.debug(inspect.currentframe().f_code.co_name)
        try:
            self.matches, self.rules = self._extract_yara_matches()
            return self.matches if self.matches else None
        except Exception as e:
            self.log.error(f"Error extracting YARA matches: {e}")
            return None

    def tag(self) -> str:
        """Return the tag for this extractor."""
        return Tag.YARA.value

    def get_clickhouse_table(self) -> str:
        """Return the ClickHouse table name for YARA matches."""
        return "yara_matches"

    def get_clickhouse_tables(self) -> Dict[str, str]:
        """Return all ClickHouse table names for multi-table support."""
        return {
            'rules': 'yara_rules',
            'matches': 'yara_matches',
        }

    def prepare_export_data(self, exporter_type: str) -> Any:
        """
        Prepare data for export to the specified exporter type.

        Uses optimized normalized schema with two tables:
        - yara_rules: Rule metadata with rule_id (UInt64 hash)
        - yara_matches: Matches with binary sha256 (32 bytes) and rule_id

        Args:
            exporter_type: The type of exporter (e.g., 'ClickHouseExporter')

        Returns:
            Data formatted for the specified exporter
        """

        if exporter_type == "ClickHouseExporter":
            if not self.matches:
                return None

            current_time = datetime.now(timezone.utc)
            source_collection = DEFAULT_SOURCE_COLLECTION

            # Convert sha256 hex to binary (32 bytes)
            sha256_binary = sha256_hex_to_binary(self.sha256)

            # Build rules data from self.rules (deduplicated by rule_id)
            rules_data = {}  # rule_id -> rule_data
            for rule_name, rule in self.rules.items():
                rule_content = f"rule {rule_name} {{ }}"  # Simplified for now
                rule_id = generate_rule_id(rule_content)
                if rule_id not in rules_data:
                    rules_data[rule_id] = [
                        rule_id,
                        rule.rule_name,
                        source_collection,
                        current_time,  # ingested_at
                        "",  # rule_text (empty for now, populated during sync)
                        json.dumps(rule.rule_meta),
                        rule.rule_tags,
                    ]

            # Build matches data
            matches_data = []
            for match in self.matches:
                rule_content = f"rule {match.rule_name} {{ }}"
                rule_id = generate_rule_id(rule_content)

                matches_data.append([
                    sha256_binary,  # Binary sha256 (32 bytes)
                    rule_id,
                    match.rule_name,
                    current_time,  # scan_date
                    match.match_strings,
                ])

            return {
                'multi_table': True,
                'rules': {
                    'table': 'yara_rules',
                    'data': list(rules_data.values()),
                    'column_names': [
                        'rule_id',
                        'rule_name',
                        'source_collection',
                        'ingested_at',
                        'rule_text',
                        'rule_meta',
                        'rule_tags',
                    ],
                    'column_type_names': [
                        'UInt64',
                        'String',
                        'LowCardinality(String)',
                        "DateTime64(3, 'UTC')",
                        'String',
                        'JSON',
                        'Array(LowCardinality(String))',
                    ],
                },
                'matches': {
                    'table': 'yara_matches',
                    'data': matches_data,
                    'column_names': [
                        'sha256',
                        'rule_id',
                        'rule_name',
                        'scan_date',
                        'match_strings',
                    ],
                    'column_type_names': [
                        'FixedString(32)',  # Binary sha256
                        'UInt64',
                        'LowCardinality(String)',
                        "DateTime64(3, 'UTC')",
                        'Array(String)',
                    ],
                },
            }

        return None


def sync_rules_to_db(source_collection: str = None) -> bool:
    """
    Sync all YARA rules from source files to the database.

    This ensures all rules exist in yara_rules table before scanning begins.
    Should be run before batch scanning or in container initialization.

    Args:
        source_collection: Source collection name (default from env)

    Returns:
        True if successful, False otherwise
    """
    from redb import settings

    source_collection = source_collection or DEFAULT_SOURCE_COLLECTION
    rules_path = YaraExtractor.get_yara_rules_path()

    if not rules_path.exists():
        print(f"[ERROR] Rules path does not exist: {rules_path}")
        return False

    # Find all rule files
    rule_files = []
    for ext in ["*.yar", "*.yara"]:
        rule_files.extend(rules_path.rglob(ext))

    if not rule_files:
        print(f"[INFO] No rule files found in {rules_path}")
        return True

    print(f"[INFO] Syncing {len(rule_files)} rule files to database...")

    current_time = datetime.now(timezone.utc)
    rules_data = []
    total_rules = 0

    for rule_file in rule_files:
        try:
            with open(rule_file, "r", encoding="utf-8") as f:
                file_content = f.read()

            # Parse individual rules from file (handles multiple rules per file)
            parsed_rules = parse_yara_rules(file_content)

            for rule in parsed_rules:
                rule_id = generate_rule_id(rule['rule_text'])

                rules_data.append([
                    rule_id,
                    rule['rule_name'],
                    source_collection,
                    current_time,
                    rule['rule_text'],
                    json.dumps(rule['rule_meta']),
                    rule['rule_tags'],
                ])
                total_rules += 1

        except Exception as e:
            print(f"[WARNING] Failed to parse rule file {rule_file}: {e}")
            continue

    print(f"[INFO] Parsed {total_rules} individual rules from {len(rule_files)} files")

    if not rules_data:
        print("[INFO] No rules to sync")
        return True

    # Insert rules to database
    try:
        client = settings.create_clickhouse_client()
        table = "yara_rules"

        client.insert(
            table,
            rules_data,
            column_names=[
                'rule_id', 'rule_name', 'source_collection',
                'ingested_at', 'rule_text', 'rule_meta', 'rule_tags'
            ],
            column_type_names=[
                'UInt64', 'String', 'LowCardinality(String)',
                "DateTime64(3, 'UTC')", 'String', 'JSON', 'Array(LowCardinality(String))'
            ]
        )

        print(f"[INFO] Successfully synced {len(rules_data)} rules to {table}")
        client.close()
        return True

    except Exception as e:
        print(f"[ERROR] Failed to sync rules to database: {e}")
        return False


# CLI utility for pre-compiling rules and syncing
if __name__ == "__main__":
    import argparse

    parser = argparse.ArgumentParser(description="YARA Rules Management Utility")
    parser.add_argument(
        "--compile",
        action="store_true",
        help="Compile all YARA rules and save to .yarac file",
    )
    parser.add_argument(
        "--sync-rules",
        action="store_true",
        help="Sync all YARA rules to the database (run before batch scanning)",
    )
    parser.add_argument(
        "--output",
        type=str,
        help="Output path for compiled rules (default: yara/compiled_rules.yarac)",
    )
    parser.add_argument(
        "--rules-path",
        type=str,
        help="Path to YARA rules directory (default: yara/)",
    )
    parser.add_argument(
        "--source-collection",
        type=str,
        help="Source collection name for rules (default: from YARA_SOURCE_COLLECTION env)",
    )

    args = parser.parse_args()

    if args.rules_path:
        os.environ["YARA_RULES_PATH"] = args.rules_path

    if args.compile:
        output_path = Path(args.output) if args.output else None
        success = YaraExtractor.compile_and_save(output_path)
        if not success:
            exit(1)

    if args.sync_rules:
        success = sync_rules_to_db(args.source_collection)
        if not success:
            exit(1)

    if not args.compile and not args.sync_rules:
        parser.print_help()