Kate R. Rosenbloom

24 papers Journal 24
YearRankTypeTitle / Venue / Authors
2022 J jnl
Nucleic Acids Res.
Brian T. Lee, Galt P. Barber, Anna Benet-Pagès, Jonathan Casper, Hiram Clawson, Mark Diekhans, Clayton M. Fischer, Jairo Navarro Gonzalez, Angie S. Hinrichs, Christopher M. Lee, Pranav Muthuraman, Luis R. Nassar, Beagan Nguy, Tiana Pereira, Gerardo Perez, Brian J. Raney, Kate R. Rosenbloom, Daniel Schmelter, Matthew L. Speir, Brittney D. Wick, Ann S. Zweig, David Haussler, Robert M. Kuhn, Maximilian Haeussler, W. James Kent
2021 J jnl
Nucleic Acids Res.
Jairo Navarro Gonzalez, Ann S. Zweig, Matthew L. Speir, Daniel Schmelter, Kate R. Rosenbloom, Brian J. Raney, Conner C. Powell, Luis R. Nassar, Nathan D. Maulding, Christopher M. Lee, Brian T. Lee, Angie S. Hinrichs, Alastair C. Fyfe, Jason D. Fernandes, Mark Diekhans, Hiram Clawson, Jonathan Casper, Anna Benet-Pagès, Galt P. Barber, David Haussler, Robert M. Kuhn, Maximilian Haeussler, W. James Kent
2020 J jnl
Nucleic Acids Res.
Christopher M. Lee, Galt P. Barber, Jonathan Casper, Hiram Clawson, Mark Diekhans, Jairo Navarro Gonzalez, Angie S. Hinrichs, Brian T. Lee, Luis R. Nassar, Conner C. Powell, Brian J. Raney, Kate R. Rosenbloom, Daniel Schmelter, Matthew L. Speir, Ann S. Zweig, David Haussler, Maximilian Haeussler, Robert M. Kuhn, W. James Kent
2019 J jnl
Nucleic Acids Res.
Maximilian Haeussler, Ann S. Zweig, Cath Tyner, Matthew L. Speir, Kate R. Rosenbloom, Brian J. Raney, Christopher M. Lee, Brian T. Lee, Angie S. Hinrichs, Jairo Navarro Gonzalez, David Gibson, Mark Diekhans, Hiram Clawson, Jonathan Casper, Galt P. Barber, David Haussler, Robert M. Kuhn, W. James Kent
2018 J jnl
Nucleic Acids Res.
Jonathan Casper, Ann S. Zweig, Chris Villarreal, Cath Tyner, Matthew L. Speir, Kate R. Rosenbloom, Brian J. Raney, Christopher M. Lee, Brian T. Lee, Donna Karolchik, Angie S. Hinrichs, Maximilian Haeussler, Luvina Guruvadoo, Jairo Navarro Gonzalez, David Gibson, Ian T. Fiddes, Christopher Eisenhart, Mark Diekhans, Hiram Clawson, Galt P. Barber, Joel Armstrong, David Haussler, Robert M. Kuhn, W. James Kent
2017 J jnl
Nucleic Acids Res.
Cath Tyner, Galt P. Barber, Jonathan Casper, Hiram Clawson, Mark Diekhans, Christopher Eisenhart, Clayton M. Fischer, David Gibson, Jairo Navarro Gonzalez, Luvina Guruvadoo, Maximilian Haeussler, Steven G. Heitner, Angie S. Hinrichs, Donna Karolchik, Brian T. Lee, Christopher M. Lee, Parisa Nejad, Brian J. Raney, Kate R. Rosenbloom, Matthew L. Speir, Chris Villarreal, John Vivian, Ann S. Zweig, David Haussler, Michael Kuhn, W. James Kent
2016 J jnl
Nucleic Acids Res.
Matthew L. Speir, Ann S. Zweig, Kate R. Rosenbloom, Brian J. Raney, Benedict Paten, Parisa Nejad, Brian T. Lee, Katrina Learned, Donna Karolchik, Angie S. Hinrichs, Steven G. Heitner, Rachel A. Harte, Maximilian Haeussler, Luvina Guruvadoo, Pauline A. Fujita, Christopher Eisenhart, Mark Diekhans, Hiram Clawson, Jonathan Casper, Galt P. Barber, David Haussler, Robert M. Kuhn, W. James Kent
2016 J jnl
Bioinform.
Angie S. Hinrichs, Brian J. Raney, Matthew L. Speir, Brooke L. Rhead, Jonathan Casper, Donna Karolchik, Robert M. Kuhn, Kate R. Rosenbloom, Ann S. Zweig, David Haussler, W. James Kent
2015 J jnl
Nucleic Acids Res.
Kate R. Rosenbloom, Joel Armstrong, Galt P. Barber, Jonathan Casper, Hiram Clawson, Mark Diekhans, Timothy R. Dreszer, Pauline A. Fujita, Luvina Guruvadoo, Maximilian Haeussler, Rachel A. Harte, Steven G. Heitner, Glenn Hickey, Angie S. Hinrichs, Robert Hubley, Donna Karolchik, Katrina Learned, Brian T. Lee, Chin H. Li, Karen H. Miga, Ngan Nguyen, Benedict Paten, Brian J. Raney, Arian F. A. Smit, Matthew L. Speir, Ann S. Zweig, David Haussler, Robert M. Kuhn, W. James Kent
2014 J jnl
Nucleic Acids Res.
Donna Karolchik, Galt P. Barber, Jonathan Casper, Hiram Clawson, Melissa S. Cline, Mark Diekhans, Timothy R. Dreszer, Pauline A. Fujita, Luvina Guruvadoo, Maximilian Haeussler, Rachel A. Harte, Steven G. Heitner, Angie S. Hinrichs, Katrina Learned, Brian T. Lee, Chin H. Li, Brian J. Raney, Brooke L. Rhead, Kate R. Rosenbloom, Cricket A. Sloan, Matthew L. Speir, Ann S. Zweig, David Haussler, Robert M. Kuhn, W. James Kent
2013 J jnl
Nucleic Acids Res.
Kate R. Rosenbloom, Cricket A. Sloan, Venkat S. Malladi, Timothy R. Dreszer, Katrina Learned, Vanessa Kirkup, Matthew C. Wong, Morgan Maddren, Ruihua Fang, Steven G. Heitner, Brian T. Lee, Galt P. Barber, Rachel A. Harte, Mark Diekhans, Jeffrey C. Long, Steven P. Wilder, Ann S. Zweig, Donna Karolchik, Robert M. Kuhn, David Haussler, W. James Kent
2013 J jnl
Nucleic Acids Res.
Laurence R. Meyer, Ann S. Zweig, Angie S. Hinrichs, Donna Karolchik, Robert M. Kuhn, Matthew C. Wong, Cricket A. Sloan, Kate R. Rosenbloom, Greg Roe, Brooke L. Rhead, Brian J. Raney, Andy Pohl, Venkat S. Malladi, Chin H. Li, Brian T. Lee, Katrina Learned, Vanessa Kirkup, Fan Hsu, Steven G. Heitner, Rachel A. Harte, Maximilian Haeussler, Luvina Guruvadoo, Mary Goldman, Belinda Giardine, Pauline A. Fujita, Timothy R. Dreszer, Mark Diekhans, Melissa S. Cline, Hiram Clawson, Galt P. Barber, David Haussler, W. James Kent
2012 J jnl
Nucleic Acids Res.
Kate R. Rosenbloom, Timothy R. Dreszer, Jeffrey C. Long, Venkat S. Malladi, Cricket A. Sloan, Brian J. Raney, Melissa S. Cline, Donna Karolchik, Galt P. Barber, Hiram Clawson, Mark Diekhans, Pauline A. Fujita, Mary Goldman, Robert C. Gravell, Rachel A. Harte, Angie S. Hinrichs, Vanessa Kirkup, Robert M. Kuhn, Katrina Learned, Morgan Maddren, Laurence R. Meyer, Andy Pohl, Brooke L. Rhead, Matthew C. Wong, Ann S. Zweig, David Haussler, W. James Kent
2012 J jnl
Nucleic Acids Res.
Timothy R. Dreszer, Donna Karolchik, Ann S. Zweig, Angie S. Hinrichs, Brian J. Raney, Robert M. Kuhn, Laurence R. Meyer, Matthew C. Wong, Cricket A. Sloan, Kate R. Rosenbloom, Greg Roe, Brooke L. Rhead, Andy Pohl, Venkat S. Malladi, Chin H. Li, Katrina Learned, Vanessa Kirkup, Fan Hsu, Rachel A. Harte, Luvina Guruvadoo, Mary Goldman, Belinda Giardine, Pauline A. Fujita, Mark Diekhans, Melissa S. Cline, Hiram Clawson, Galt P. Barber, David Haussler, W. James Kent
2011 J jnl
Nucleic Acids Res.
Brian J. Raney, Melissa S. Cline, Kate R. Rosenbloom, Timothy R. Dreszer, Katrina Learned, Galt P. Barber, Laurence R. Meyer, Cricket A. Sloan, Venkat S. Malladi, Krishna M. Roskin, Bernard B. Suh, Angie S. Hinrichs, Hiram Clawson, Ann S. Zweig, Vanessa Kirkup, Pauline A. Fujita, Brooke L. Rhead, Kayla E. Smith, Andy Pohl, Robert M. Kuhn, Donna Karolchik, David Haussler, W. James Kent
2011 J jnl
Nucleic Acids Res.
Pauline A. Fujita, Brooke L. Rhead, Ann S. Zweig, Angie S. Hinrichs, Donna Karolchik, Melissa S. Cline, Mary Goldman, Galt P. Barber, Hiram Clawson, António Coelho, Mark Diekhans, Timothy R. Dreszer, Belinda Giardine, Rachel A. Harte, Jennifer Hillman-Jackson, Fan Hsu, Vanessa Kirkup, Robert M. Kuhn, Katrina Learned, Chin H. Li, Laurence R. Meyer, Andy Pohl, Brian J. Raney, Kate R. Rosenbloom, Kayla E. Smith, David Haussler, W. James Kent
2010 J jnl
Nucleic Acids Res.
Kate R. Rosenbloom, Timothy R. Dreszer, Michael Pheasant, Galt P. Barber, Laurence R. Meyer, Andy Pohl, Brian J. Raney, Ting Wang, Angie S. Hinrichs, Ann S. Zweig, Pauline A. Fujita, Katrina Learned, Brooke L. Rhead, Kayla E. Smith, Robert M. Kuhn, Donna Karolchik, David Haussler, W. James Kent
2010 J jnl
Nucleic Acids Res.
Brooke L. Rhead, Donna Karolchik, Robert M. Kuhn, Angie S. Hinrichs, Ann S. Zweig, Pauline A. Fujita, Mark Diekhans, Kayla E. Smith, Kate R. Rosenbloom, Brian J. Raney, Andy Pohl, Michael Pheasant, Laurence R. Meyer, Katrina Learned, Fan Hsu, Jennifer Hillman-Jackson, Rachel A. Harte, Belinda Giardine, Timothy R. Dreszer, Hiram Clawson, Galt P. Barber, David Haussler, W. James Kent
2009 J jnl
Nucleic Acids Res.
Robert M. Kuhn, Donna Karolchik, Ann S. Zweig, Ting Wang, Kayla E. Smith, Kate R. Rosenbloom, Brooke L. Rhead, Brian J. Raney, Andy Pohl, Michael Pheasant, Laurence R. Meyer, Fan Hsu, Angela S. Hinrichs, Rachel A. Harte, Belinda Giardine, Pauline A. Fujita, Mark Diekhans, Timothy R. Dreszer, Hiram Clawson, Galt P. Barber, David Haussler, W. James Kent
2008 J jnl
Nucleic Acids Res.
Donna Karolchik, Robert M. Kuhn, Robert Baertsch, Galt P. Barber, Hiram Clawson, Mark Diekhans, Belinda Giardine, Rachel A. Harte, Angela S. Hinrichs, Fan Hsu, Kord M. Kober, Webb Miller, Jakob Skou Pedersen, Andy Pohl, Brian J. Raney, Brooke L. Rhead, Kate R. Rosenbloom, Kayla E. Smith, Mario Stanke, Archana Thakkapallayil, Heather Trumbower, Ting Wang, Ann S. Zweig, David Haussler, W. James Kent
2007 J jnl
Nucleic Acids Res.
Daryl J. Thomas, Kate R. Rosenbloom, Hiram Clawson, Angie S. Hinrichs, Heather Trumbower, Brian J. Raney, Donna Karolchik, Galt P. Barber, Rachel A. Harte, Jennifer Hillman-Jackson, Robert M. Kuhn, Brooke L. Rhead, Kayla E. Smith, Archana Thakkapallayil, Ann S. Zweig, David Haussler, W. James Kent
2007 J jnl
Nucleic Acids Res.
Robert M. Kuhn, Donna Karolchik, Ann S. Zweig, Heather Trumbower, Daryl J. Thomas, Archana Thakkapallayil, Charles W. Sugnet, Mario Stanke, Kayla E. Smith, Adam C. Siepel, Kate R. Rosenbloom, Brooke L. Rhead, Brian J. Raney, Andy Pohl, Jakob Skou Pedersen, Fan Hsu, Angela S. Hinrichs, Rachel A. Harte, Mark Diekhans, Hiram Clawson, Gill Bejerano, Galt P. Barber, Robert Baertsch, David Haussler, W. James Kent
2006 J jnl
PLoS Comput. Biol.
Jakob Skou Pedersen, Gill Bejerano, Adam C. Siepel, Kate R. Rosenbloom, Kerstin Lindblad-Toh, Eric S. Lander, Jim Kent, Webb Miller, David Haussler
2006 J jnl
Nucleic Acids Res.
Angela S. Hinrichs, Donna Karolchik, Robert Baertsch, Galt P. Barber, Gill Bejerano, Hiram Clawson, Mark Diekhans, Terrence S. Furey, Rachel A. Harte, Fan Hsu, Jennifer Hillman-Jackson, Robert M. Kuhn, Jakob Skou Pedersen, Andy Pohl, Brian J. Raney, Kate R. Rosenbloom, Adam C. Siepel, Kayla E. Smith, Charles W. Sugnet, A. Sultan-Qurraie, Daryl J. Thomas, Heather Trumbower, Ryan J. Weber, M. Weirauch, Ann S. Zweig, David Haussler, W. James Kent
redb/extractors/decompiler/_archive/DecompileGhidra.py
← Index redb/extractors/decompiler/_archive/DecompileGhidra.py python
from hashlib import sha256, md5
import inspect
import subprocess
import json
import os
import time
from datetime import datetime, timezone
from typing import Dict, List, Any, Optional

from dotenv import load_dotenv
from redb.extractors.enum import Tag
from redb.extractors.extractor import Extractor
import magic
import pefile
import ppdeep
import tlsh


class DecompileGhidra(Extractor):
    def __init__(
        self,
        filepath,
        log,
        exporters=None,
        index_prefix=None,
        elastic_index=None,
        known_benign=False,
        known_malicious=False,
        filetype=None,
    ):
        super().__init__(
            filepath,
            log,
            exporters,
            index_prefix,
            elastic_index,
            known_benign,
            known_malicious,
        )
        self.log.debug(inspect.currentframe().f_code.co_name)
        self.ghidra_path = os.getenv("GHIDRA_PATH", "/opt/ghidra")
        self.java_script_path = os.getenv(
            "GHIDRA_SCRIPT_PATH",
            "/opt/ghidra/Ghidra/Features/Base/ghidra_scripts/GhidraDecompilerScript.java",
        )
        self.analysis_results = None
        self.ghidra_process = None  # Track the current process
        self.project_path = None
        self.filetype = filetype

        # Convert TIMEOUT to integer with a default of 1200 seconds (20 minutes)
        try:
            self.TIMEOUT = int(os.getenv("GHIDRA_TIMEOUT", "1200"))
        except ValueError:
            self.log.warning(
                "Invalid GHIDRA_TIMEOUT value, using default of 1200 seconds"
            )
            self.TIMEOUT = 1200

        self.initialize_project()

    def __enter__(self):
        return self

    def __exit__(self, exc_type, exc_val, exc_tb):
        self.cleanup_run()

    def is_dotnet(self):
        try:
            if self.filetype == "pebin":
                file_type = magic.from_buffer(self.binary)
                if ".Net" in file_type:
                    return True
                pe = pefile.PE(self.filepath)
                for entry in pe.OPTIONAL_HEADER.DATA_DIRECTORY:
                    # IMAGE_DIRECTORY_ENTRY_COM_DESCRIPTOR is typically 14
                    if (
                        entry.name == "IMAGE_DIRECTORY_ENTRY_COM_DESCRIPTOR"
                        and entry.Size > 0
                    ):
                        return True
                return False
        except AttributeError as e:
            self.log.error(
                f"AttributeError error dotnet file {self.hash.sha256} Full error : {e}"
            )
            return False

    def cleanup_run(self):
        """Clean up after analysis."""
        try:
            if self.ghidra_process and self.ghidra_process.poll() is None:
                self.ghidra_process.terminate()
                try:
                    self.ghidra_process.wait(timeout=5)
                except subprocess.TimeoutExpired:
                    self.ghidra_process.kill()

            # Clean up project directory
            if self.project_path and os.path.exists(self.project_path):
                import shutil

                shutil.rmtree(self.project_path)
                self.log.debug(f"Cleaned up project directory: {self.project_path}")

            # Force garbage collection
            import gc

            gc.collect()
        except Exception as e:
            self.log.error(f"Error in cleanup: {e}")

    # @classmethod
    # def cleanup_batch(cls):
    #     """Clean up the persistent project at the end of a batch."""
    #     print(f"Cleaning up Ghidra project for batch")
    #     if cls._project_path and os.path.exists(cls._project_path):
    #         try:
    #             import shutil
    #             shutil.rmtree(cls._project_path)
    #             cls._project_initialized = False
    #             cls._project_path = None
    #         except Exception as e:
    #             print(f"Error cleaning up project: {e}")

    def _get_environment(self):
        """Setup and return the environment for Ghidra."""
        env = os.environ.copy()
        java_home = os.getenv("GHIDRA_JAVA_HOME", "/usr/lib/jvm/java-17-openjdk-amd64")
        env.update(
            {
                "JAVA_HOME": java_home,
                "PATH": f"{java_home}/bin:{env['PATH']}",
                "LD_LIBRARY_PATH": f"{java_home}/lib:{env.get('LD_LIBRARY_PATH', '')}",
            }
        )
        # Print environment variables for debugging
        self.log.debug(f"JAVA_HOME: {env['JAVA_HOME']}")
        self.log.debug(f"PATH: {env['PATH']}")
        self.log.debug(f"LD_LIBRARY_PATH: {env['LD_LIBRARY_PATH']}")

        return env

    def initialize_project(self):
        """Initialize a temporary Ghidra project for this file."""
        # Create unique project directory
        self.project_path = f"/tmp/ghidra_{os.path.basename(self.filepath)}_{str(int(time.time()))}_{os.getpid()}"
        os.makedirs(self.project_path, exist_ok=True)
        self.log.debug(f"Created temporary project at {self.project_path}")

        # Create a minimal initialization file
        init_file = os.path.join(self.project_path, ".init")
        with open(init_file, "wb") as f:
            f.write(bytes([0x7F, 0x45, 0x4C, 0x46]))  # Valid ELF header magic bytes

        # Initialize project with minimal file
        env = self._get_environment()
        cmd = [
            f"{self.ghidra_path}/support/analyzeHeadless",
            self.project_path,
            "TempProject",
            "-import",
            init_file,
        ]

        try:
            result = subprocess.run(cmd, env=env, capture_output=True, text=True)
            if result.returncode != 0:
                self.log.error(f"Failed to initialize project: {result.stderr}")
                raise RuntimeError("Project initialization failed")

            # Clean up initialization file
            os.remove(init_file)
            self.log.debug("Project initialized successfully")

        except Exception as e:
            self.log.error(f"Error initializing project: {e}")
            raise

    def analyze_binary(self) -> Optional[Dict[str, Any]]:
        """Run Ghidra analysis and return results."""
        self.log.debug("Starting binary analysis")

        # # Check if packed
        # if self.check_binary_protection():
        #     self.log.warning("Skipping protected binary")
        #     return None

        # Check for .NET only if needed
        # if self.is_dotnet():
        #     self.MAX_NAMED_ARG_WARNINGS = 10000  # Higher threshold for .NET
        #     self.log.info("Adjusting parameters for .NET binary")
        # else:
        #     self.MAX_NAMED_ARG_WARNINGS = 1000  # Normal threshold

        if not os.path.exists(self.java_script_path):
            self.log.error(f"Java script not found: {self.java_script_path}")
            return None

        env = self._get_environment()

        try:
            base_cmd = [
                f"{self.ghidra_path}/support/analyzeHeadless",
                self.project_path,
                "TempProject",
                "-import",
                self.filepath,
                "-scriptPath",
                os.path.dirname(self.java_script_path),
                "-postScript",
                self.java_script_path,
                self.sha256,
                self.filepath,
            ]
            return self.run_ghidra(base_cmd, env)

        except Exception as e:
            self.log.error(f"Error in Ghidra analysis: {e}")
            return None

        finally:
            self.cleanup_run()

    def run_ghidra(
        self, cmd: list, env: Optional[Dict[str, str]] = None
    ) -> Optional[Dict[str, Any]]:
        """Run Ghidra process and capture JSON output with improved logging separation."""
        process = None
        try:
            self.log.info(f"Starting Ghidra analysis: {' '.join(cmd)}")
            start_time = time.time()

            process = subprocess.Popen(
                cmd, env=env, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True
            )
            self.ghidra_process = process

            warning_counter = 0
            named_arg_counter = 0
            # Read all output lines
            json_output = None
            while True:
                line = process.stdout.readline()
                if not line and process.poll() is not None:
                    break

                stripped_line = line.strip()
                if not stripped_line:
                    continue

                # if 'Invalid FieldOrProp value in NamedArg' in stripped_line:
                #     named_arg_counter += 1
                #     if named_arg_counter > self.MAX_NAMED_ARG_WARNINGS:
                #         self.log.error(f"Too many NamedArg warnings ({named_arg_counter}), possible protected file.")
                #         self.ghidra_process.kill()
                #         return None
                if (
                    stripped_line.startswith("{")
                    and '"sha256"' in stripped_line
                    and '"decompiled"' in stripped_line
                ):
                    # This is our actual JSON output from GhidraDecompilerScript
                    json_output = stripped_line
                elif any(level in stripped_line for level in ["INFO", "WARN", "ERROR"]):
                    # Ghidra framework logging
                    log_level = (
                        "debug"
                        if "INFO" in stripped_line
                        else "warning"
                        if "WARN" in stripped_line
                        else "error"
                    )
                    if log_level == "warning" and any(
                        expected in stripped_line
                        for expected in [
                            "Unable to disassemble EXTERNAL block",
                            "Failed to markup ELF Note",
                            "Invalid FieldOrProp value in NamedArg",
                            "Unable to resolve constructor",
                            "Could not follow disassembly flow into non-existing memory",
                            "Unable to read bytes at ram",
                        ]
                    ):
                        # Skip expected warnings
                        continue

                    getattr(self.log, log_level)(f"Ghidra info: {stripped_line}")

            # Process completion and stderr
            try:
                stderr = process.stderr.read()
                process.wait(timeout=self.TIMEOUT)

                if stderr:
                    for line in stderr.splitlines():
                        stripped_line = line.strip()
                        if not stripped_line:
                            continue
                        if "ERROR" in stripped_line:
                            self.log.error(f"Ghidra stderr: {stripped_line}")
                        elif "WARN" in stripped_line:
                            self.log.warning(f"Ghidra stderr: {stripped_line}")
                        else:
                            self.log.debug(f"Ghidra stderr: {stripped_line}")

            except subprocess.TimeoutExpired:
                process.kill()
                self.log.error("Ghidra analysis timed out")
                return None

            elapsed_time = time.time() - start_time
            self.log.debug(f"Ghidra analysis completed in {elapsed_time:.2f}s")

            # Parse JSON output if we found it
            if json_output:
                try:
                    result = json.loads(json_output)
                    # Validate the required structure
                    if not isinstance(result, dict) or not all(
                        k in result
                        for k in ["sha256", "decompiled", "disassembled", "cfg"]
                    ):
                        self.log.error("Invalid JSON structure from Ghidra")
                        return None
                    return result
                except json.JSONDecodeError as e:
                    self.log.error(f"Failed to parse Ghidra JSON output: {e}")
                    return None
            else:
                self.log.error("No JSON output received from Ghidra")
                return None

        except Exception as e:
            self.log.error(f"Error running Ghidra: {str(e)}")
            if hasattr(e, "__traceback__"):
                import traceback

                self.log.debug(
                    f"Traceback: {''.join(traceback.format_tb(e.__traceback__))}"
                )
            return None

        finally:
            if process:
                try:
                    # Ensure pipes are closed
                    if process.stdout:
                        process.stdout.close()
                    if process.stderr:
                        process.stderr.close()
                    # Terminate process if still running
                    if process.poll() is None:
                        process.terminate()
                        try:
                            process.wait(timeout=5)
                        except subprocess.TimeoutExpired:
                            process.kill()
                except Exception as e:
                    self.log.error(f"Error cleaning up Ghidra process: {e}")

    def extract(self) -> bool:
        """Extract and process all analysis results."""
        self.log.debug(inspect.currentframe().f_code.co_name)
        try:
            results = self.analyze_binary()
            if not results:
                return False

            self.analysis_results = results
            return True

        except Exception as e:
            self.log.error(f"Error in extraction: {e}")
            return False

    def prepare_export_data(self, exporter_type: str) -> Any:
        """Prepare data for database export."""
        self.log.debug(inspect.currentframe().f_code.co_name)
        if not self.analysis_results:
            return None

        if exporter_type == "ClickHouseExporter":
            now = datetime.now(timezone.utc)

            def prepare_array_field(value, array_type):
                """Helper to prepare array fields with proper null handling"""
                if value is None:
                    return []
                return value

            def ssdeep_disassembly(func):
                try:
                    if len(func) > 1:
                        return ppdeep.hash(func)
                    return ""
                except Exception as e:
                    self.log.error(f"Error in disassembly ssdeep hash calculation: {e}")
                    return ""

            def tlsh_disassembly(func):
                try:
                    if len(func) >= 50:
                        return tlsh.hash(func.encode("utf-8"))
                    return ""
                except Exception as e:
                    self.log.error(f"Error in disassembly tlsh hash calculation: {e}")
                    return ""

            return {
                "multi_table": True,
                "decompiled_content": {
                    "table": "decompiled_functions_content",
                    "data": [
                        [
                            f["decompiled_content_hash"],
                            f["decompiled_function"],
                            f["function_type"],
                            now,
                        ]
                        for f in self.analysis_results["decompiled"]
                    ],
                    "column_names": [
                        "decompiled_content_hash",
                        "decompiled_function",
                        "function_type",
                        "analysis_date",
                    ],
                    "column_type_names": [
                        "FixedString(64)",
                        "String",
                        "Enum8('USER'=1, 'LIBRARY'=2, 'THUNK'=3, 'EXTERNAL'=4, 'UNKNOWN'=5)",
                        "DateTime64(3, 'UTC')",
                    ],
                },
                "decompiled_refs": {
                    "table": "decompiled_functions_references",
                    "data": [
                        [
                            self.analysis_results["sha256"],
                            f["decompiled_content_hash"],
                            f["decompiled_function_name"],
                            f["decompiled_function_address"],
                            now,
                        ]
                        for f in self.analysis_results["decompiled"]
                    ],
                    "column_names": [
                        "sha256",
                        "decompiled_content_hash",
                        "decompiled_function_name",
                        "decompiled_function_address",
                        "analysis_date",
                    ],
                    "column_type_names": [
                        "FixedString(64)",
                        "FixedString(64)",
                        "LowCardinality(String)",
                        "String",
                        "DateTime64(3, 'UTC')",
                    ],
                },
                "disassembled_content": {
                    "table": "disassembled_functions_content",
                    "data": [
                        [
                            f["disassembled_content_hash"],
                            f["fully_normalized_content_hash"],
                            f["api_normalized_content_hash"],
                            f["category_normalized_content_hash"],
                            f.get("disassembled_function", ""),
                            f.get("fully_normalized_disassembly", ""),
                            f.get("api_normalized_disassembly", ""),
                            f.get("category_normalized_disassembly", ""),
                            ssdeep_disassembly(f.get("disassembled_function", "")),
                            tlsh_disassembly(f.get("disassembled_function", "")),
                            ssdeep_disassembly(
                                f.get("fully_normalized_disassembly", "")
                            ),
                            tlsh_disassembly(f.get("fully_normalized_disassembly", "")),
                            f.get("function_type", "UNKNOWN"),
                            f.get("instruction_count", 0),
                            prepare_array_field(
                                f.get("instruction_types"), "LowCardinality(String)"
                            ),
                            f.get("control_flow_count", 0),
                            prepare_array_field(
                                f.get("memory_access_pattern"), "LowCardinality(String)"
                            ),
                            prepare_array_field(
                                f.get("register_usage"), "LowCardinality(String)"
                            ),
                            f.get("data_references_count", 0),
                            # prepare_array_field(f.get('opcode_frequency_vector'), 'Float32'),
                            # prepare_array_field(f.get('api_calls_vector'), 'Float32'),
                            # prepare_array_field(f.get('minhash_signature'), 'UInt64'),
                            # f.get('pic_hash', ''),
                            f.get("max_block_size", 0),
                            f.get("num_calls", 0),
                            f.get("stack_size", 0),
                            # prepare_array_field(f.get('instruction_type_ratios'), 'Float32'),
                            # prepare_array_field(f.get('instruction_embedding'), 'Float32'),
                            now,
                        ]
                        for f in self.analysis_results["disassembled"]
                    ],
                    "column_names": [
                        "disassembled_content_hash",
                        "fully_normalized_content_hash",
                        "api_normalized_content_hash",
                        "category_normalized_content_hash",
                        "disassembled_function",
                        "fully_normalized_disassembly",
                        "api_normalized_disassembly",
                        "category_normalized_disassembly",
                        "ssdeep_disassembly",
                        "tlsh_disassembly",
                        "ssdeep_fully_normalized",
                        "tlsh_fully_normalized",
                        "function_type",
                        "instruction_count",
                        "instruction_types",
                        "control_flow_count",
                        "memory_access_pattern",
                        "register_usage",
                        "data_references_count",
                        # 'opcode_frequency_vector',
                        # 'api_calls_vector',
                        # 'minhash_signature',
                        # 'pic_hash',
                        "max_block_size",
                        "num_calls",
                        "stack_size",
                        # 'instruction_type_ratios',
                        # 'instruction_embedding',
                        "analysis_date",
                    ],
                    "column_type_names": [
                        "FixedString(64)",
                        "FixedString(64)",
                        "FixedString(64)",
                        "FixedString(64)",
                        "String",
                        "String",
                        "String",
                        "String",
                        "Nullable(String)",
                        "Nullable(FixedString(72))",
                        "Nullable(String)",
                        "Nullable(FixedString(72))",
                        "Enum8('USER'=1, 'LIBRARY'=2, 'THUNK'=3, 'EXTERNAL'=4, 'UNKNOWN'=5)",
                        "UInt32",
                        "Array(LowCardinality(String))",
                        "UInt32",
                        "Array(LowCardinality(String))",
                        "Array(LowCardinality(String))",
                        "UInt32",
                        # 'Array(Float32)',
                        # 'Array(Float32)',
                        # 'Array(UInt64)',
                        # 'Nullable(FixedString(16))',
                        "Nullable(UInt32)",
                        "Nullable(UInt32)",
                        "Nullable(Int32)",
                        # 'Array(Float32)',
                        # 'Array(Float32)',
                        "DateTime64(3, 'UTC')",
                    ],
                },
                "disassembled_refs": {
                    "table": "disassembled_functions_references",
                    "data": [
                        [
                            self.analysis_results["sha256"],
                            f["disassembled_content_hash"],
                            f["fully_normalized_content_hash"],
                            f["api_normalized_content_hash"],
                            f["category_normalized_content_hash"],
                            f["disassembled_function_name"],
                            f["disassembled_function_address"],
                            now,
                        ]
                        for f in self.analysis_results["disassembled"]
                    ],
                    "column_names": [
                        "sha256",
                        "disassembled_content_hash",
                        "fully_normalized_content_hash",
                        "api_normalized_content_hash",
                        "category_normalized_content_hash",
                        "disassembled_function_name",
                        "disassembled_function_address",
                        "analysis_date",
                    ],
                    "column_type_names": [
                        "FixedString(64)",
                        "FixedString(64)",
                        "FixedString(64)",
                        "FixedString(64)",
                        "FixedString(64)",
                        "LowCardinality(String)",
                        "String",
                        "DateTime64(3, 'UTC')",
                    ],
                },
                "cfg_blocks": {
                    "table": "cfg_blocks",
                    "data": [
                        [
                            b["block_id"],
                            self.analysis_results["sha256"],
                            b["function_address"],
                            b["block_start_address"],
                            b["block_end_address"],
                            b["block_size"],
                            b["block_instructions"],
                            b["fully_normalized_instructions"],
                            b["api_normalized_instructions"],
                            b["category_normalized_instructions"],
                            b.get("predecessor_blocks", []),
                            b.get("successor_blocks", []),  # Use empty array as default
                            b.get("is_entry_block", False),
                            b.get("is_exit_block", False),
                            b.get("branch_type", "UNKNOWN"),
                            b.get("referenced_constants", []),
                            b.get("sign", 1),  # Use 1 as default for sign
                            now,
                        ]
                        for b in self.analysis_results["cfg"]
                    ],
                    "column_names": [
                        "block_id",
                        "sha256",
                        "function_address",
                        "block_start_address",
                        "block_end_address",
                        "block_size",
                        "block_instructions",
                        "fully_normalized_instructions",
                        "api_normalized_instructions",
                        "category_normalized_instructions",
                        "predecessor_blocks",
                        "successor_blocks",
                        "is_entry_block",
                        "is_exit_block",
                        "branch_type",
                        "referenced_constants",
                        "sign",
                        "analysis_date",
                    ],
                    "column_type_names": [
                        "FixedString(64)",
                        "FixedString(64)",
                        "String",
                        "String",
                        "String",
                        "UInt32",
                        "String",
                        "Nullable(String)",
                        "Nullable(String)",
                        "Nullable(String)",
                        "Array(String)",
                        "Array(String)",
                        "Bool",
                        "Bool",
                        "Enum8('DIRECT'=1, 'CONDITIONAL'=2, 'CALL'=3, 'RETURN'=4, 'FALLTHROUGH'=5, 'UNKNOWN'=6)",
                        "Array(String)",
                        "Int8",
                        "DateTime64(3, 'UTC')",
                    ],
                },
                "function_analysis_errors": {
                    "table": "function_analysis_errors",
                    "data": [
                        [
                            self.analysis_results["sha256"],
                            f["function_name"],
                            f["function_address"],
                            f["error_location"],
                            f.get(
                                "error_message", ""
                            ),  # it could be empty, how to handle it?
                            f.get("error_details", ""),
                            f.get("error_type", "unknown"),
                            md5(
                                f"{f['error_message']}{f['function_name']}{f['function_address']}{f['error_location']}".encode()
                            ).hexdigest(),
                            "new",
                            now,
                        ]
                        for f in self.analysis_results["errors"]
                    ],
                    "column_names": [
                        "sha256",
                        "function_name",
                        "function_address",
                        "error_location",
                        "error_message",
                        "error_details",
                        "error_type",
                        "error_hash",
                        "status",
                        "analysis_date",
                    ],
                    "column_type_names": [
                        "FixedString(64)",
                        "Nullable(String)",
                        "String",
                        "LowCardinality(String)",
                        "Nullable(String)",
                        "Nullable(String)",
                        "Nullable(String)",
                        "FixedString(32)",
                        "Enum8('new'=1, 'investigating'=2, 'fixed'=3, 'wontfix'=4)",
                        "DateTime64(3, 'UTC')",
                    ],
                },
            }

    def tag(self) -> str:
        """Return the tag for this extractor."""
        return Tag.DECOMPILED.value

    def get_clickhouse_table(self) -> str:
        """Not used directly as we're handling multiple tables."""
        pass