Katherine K. Kim

34 papers Misc 12Journal 17Unranked 5
YearRankTypeTitle / Venue / Authors
2026 J jnl
CoRR
Scott P. McGrath, Katherine K. Kim, Karnjit Johl, Haibo Wang, Nick Anderson
2025 J jnl
J. Am. Medical Informatics Assoc.
Katherine K. Kim, Uba Backonja
2025 J jnl
CoRR
Katherine K. Kim, Scott McGrath, David Lindeman
2025 conf
MedInfo
Scott P. McGrath, Katherine K. Kim
2024 J jnl
J. Am. Medical Informatics Assoc.
Katherine K. Kim, Uba Backonja
2023 J jnl
J. Am. Medical Informatics Assoc.
Tsung-Ting Kuo, Anh Pham, Maxim E. Edelson, Jihoon Kim, Jason Chan, Yash Gupta, Lucila Ohno-Machado, David M. Anderson, Chandrasekar Balacha, Tyler Bath, Sally L. Baxter, Andrea Becker-Pennrich, Douglas S. Bell, Elmer V. Bernstam, Ngan Chau, Michele E. Day, Jason N. Doctor, Scott L. DuVall, Robert El-Kareh, Renato Florian, Robert W. Follett, Benjamin P. Geisler, Alessandro Ghigi, Assaf Gottlieb, Ludwig Christian G. Hinske, Zhaoxian Hu, Diana Ir, Xiaoqian Jiang, Katherine K. Kim, Tara K. Knight, Jejo D. Koola, Nelson Lee, Ulrich Mansmann, Michael E. Matheny, Daniella Meeker, Zongyang Mou, Larissa Neumann, Nghia H. Nguyen, Nick Anderson, Eunice Park, Paulina Paul, Mark J. Pletcher, Kai W. Post, Clemens Rieder, Clemens Scherer, Lisa M. Schilling, Andrey Soares, Spencer L. SooHoo, Ekin Soysal, Steven Covington, Brian Tep, Brian Toy, Baocheng Wang, Zhen R. Wu, Hua Xu, Yong K. Choi, Kai Zheng, Yujia Zhou, Rachel A Zucker
2023 conf
MedInfo
Sayantani Sarkar, Katherine K. Kim, Xin Liu, Jill G. Joseph, Joanne Natale
2023 J jnl
J. Am. Medical Informatics Assoc.
Brad Morse, Katherine K. Kim, Zixuan Xu, Cynthia G. Matsumoto, Lisa M. Schilling, Lucila Ohno-Machado, Selene S. Mak, Michelle S. Keller
2022 Misc conf
AMIA
Sayantani Sarkar, Katherine K. Kim
2022 Misc conf
AMIA
Melissa A. Bruno, Scott McGrath, Katherine K. Kim
2021 Misc conf
AMIA
Scott P. McGrath, Cynthia G. Matsumoto, David Lindeman, Katherine K. Kim
2021 J jnl
J. Am. Medical Informatics Assoc.
Rupa S. Valdez, Don E. Detmer, Philip E. Bourne, Katherine K. Kim, Robin Austin, Anna McCollister-Slipp, Courtney C. Rogers, Karen C. Waters-Wicks
2021 J jnl
J. Am. Medical Informatics Assoc.
Jihoon Kim, Larissa Neumann, Paulina Paul, Michele E. Day, Michael Aratow, Douglas S. Bell, Jason N. Doctor, Ludwig Christian G. Hinske, Xiaoqian Jiang, Katherine K. Kim, Michael E. Matheny, Daniella Meeker, Mark J. Pletcher, Lisa M. Schilling, Spencer L. SooHoo, Hua Xu, Kai Zheng, Lucila Ohno-Machado
2020 J jnl
CoRR
Yiran Li, Takanori Fujiwara, Yong K. Choi, Katherine K. Kim, Kwan-Liu Ma
2020 J jnl
Vis. Informatics
Yiran Li, Takanori Fujiwara, Yong K. Choi, Katherine K. Kim, Kwan-Liu Ma
2020 Misc conf
AMIA
Victoria Ngo, Theresa H. Keegan, Brian A. Jonas, Mike Hogarth, Katherine K. Kim
2020 J jnl
Health Informatics J.
Mohammad Reza Naeemabadi, Jesper Hessellund Søndergaard, Anita Klastrup, Anne Philbert Schlünsen, Rikke Emilie Kildahl Lauritsen, John Hansen, Niels Kragh Madsen, Ole Simonsen, Ole Kæseler Andersen, Katherine K. Kim, Birthe Dinesen
2019 Misc conf
AMIA
Yong Kyung Choi, Javier E. Lopez, Daniella Meeker, Lucila Ohno-Machado, Katherine K. Kim
2017 Misc conf
AMIA
Sarah C. Haynes, Katherine K. Kim
2017 J jnl
J. Am. Medical Informatics Assoc.
Dmitry Khodyakov, Sean Grant, Daniella Meeker, Marika Booth, Nathaly Pacheco-Santivanez, Katherine K. Kim
2017 Misc conf
AMIA
Katherine K. Kim, Lucila Ohno-Machado
2016 conf
Nursing Informatics
Sarah C. Haynes, Katherine K. Kim
2016 conf
Nursing Informatics
Katherine K. Kim, Janice F. Bell, Richard Bold, Andra Davis, Victoria Ngo, Sarah C. Reed, Jill G. Joseph
2016 Misc conf
AMIA
Katherine K. Kim, Dmitry Khodyakov, Kate Marie, Marika Booth, Paul A. Heidenreich, Zhaoping Li, Michael K. Ong, Jane C. Burns, Daniella Meeker, Lucila Ohno-Machado
2016 Misc conf
AMIA
Pei-Yun Sabrina Hsueh, Susan Peterson, Fernando José Martín-Sánchez, Katherine K. Kim, Çagatay Demiralp
2015 Misc conf
AMIA
Katherine K. Kim, Janice Bell, Charles Boicey, Janet Freeman-Daily, Anna McCollister-Slipp, Jill G. Joseph
2015 J jnl
J. Am. Medical Informatics Assoc.
Daniella Meeker, Xiaoqian Jiang, Michael E. Matheny, Claudiu Farcas, Michel D'Arcy, Laura Pearlman, Lavanya Nookala, Michele E. Day, Katherine K. Kim, Hyeoneui Kim, Aziz A. Boxwala, Robert E. El-Kareh, Grace Kuo, Frederic S. Resnic, Carl Kesselman, Lucila Ohno-Machado
2015 J jnl
J. Am. Medical Informatics Assoc.
Katherine K. Kim, Jill G. Joseph, Lucila Ohno-Machado
2015 J jnl
Pers. Ubiquitous Comput.
Katherine K. Kim, Holly C. Logan, Edmund Young, Christina M. Sabee
2014 Misc conf
AMIA
Katherine K. Kim, Brian Goodness, Holly C. Logan, David A. Minch, Dolores Yanagihara
2014 conf
CTS
Katherine K. Kim, Janice Bell, Sarah C. Reed, Jill G. Joseph, Richard Bold, Kimberlie L. Cerrone, Daniel Altobello, Joydip Homchowdhury
2014 J jnl
J. Am. Medical Informatics Assoc.
Lucila Ohno-Machado, Zia Agha, Douglas S. Bell, Lisa Dahm, Michele E. Day, Jason N. Doctor, Davera Gabriel, Maninder K. Kahlon, Katherine K. Kim, Michael A. Hogarth, Michael E. Matheny, Daniella Meeker, Jonathan R. Nebeker
2014 J jnl
J. Am. Medical Informatics Assoc.
Katherine K. Kim, Dennis K. Browe, Holly C. Logan, Roberta Holm, Lori Hack, Lucila Ohno-Machado
2013 Misc conf
AMIA
Laura Mamo, Dennis K. Browe, Holly M. Logan, Katherine K. Kim
redb/extractors/ioc_extractor/ioc_extractor.py
← Index redb/extractors/ioc_extractor/ioc_extractor.py python
"""
IOC Extractor - Extractor class for extracting IOCs from decompilation results.

This extractor works with in-memory data from DecompileBinja, following the
standard Extractor pattern to support both ClickHouse and PrintExporter (dry-run).

Usage:
    # After DecompileBinja completes:
    ioc_extractor = IOCExtractorFromResults(
        analysis_results=decompiler.analysis_results,
        sha256=sha256,
        log=logger,
        exporters=exporters,
        index_prefix=index_prefix
    )
    ioc_extractor.export_data()
"""

import inspect
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, List, Dict, Optional

from redb.extractors.enum import Tag
from redb.extractors.database_exporters import DatabaseExporter

# Import the IOCScraper and related classes from standalone module
from redb.extractors.ioc_extractor.standalone_ioc_extractor import (
    IOCScraper,
    IOCType,
    SourceType,
    ExtractedIOC,
)
from typing import Set


class IOCExtractorFromResults:
    """
    Extracts IOCs from in-memory decompilation results.

    This follows a simplified Extractor pattern but doesn't inherit from Extractor
    since it doesn't read from a binary file - instead it takes already-processed
    analysis results from DecompileBinja.
    """

    def __init__(
        self,
        analysis_results: Dict[str, Any],
        sha256: str,
        log: Any,
        exporters: Optional[List[DatabaseExporter]] = None,
        index_prefix: Optional[str] = None,
        tld_file: Optional[Path] = None,
        suppress_types: Optional[Set[IOCType]] = None,
        js_context: bool = False,
    ):
        """
        Initialize IOC Extractor with analysis results.

        Args:
            analysis_results: Dict containing 'strings' and 'decompiled' lists from DecompileBinja
            sha256: Sample SHA256 hash
            log: Logger instance
            exporters: List of database exporters (ClickHouse, Print, etc.)
            index_prefix: Index prefix for database
            tld_file: Optional path to TLD list file
            js_context: When True, the underlying IOCScraper rejects FQDN
                candidates that match JS object-access syntax (see
                JS_FP_TLDS / JS_FP_SLDS). Set this for the JS pipeline only;
                APK suppresses FQDN entirely via suppress_types and binary
                callers leave it disabled.
        """
        self.log = log
        self.log.debug(f"Creating {self.__class__.__name__}")
        self.analysis_results = analysis_results
        self.sha256 = sha256
        self.exporters = exporters or []
        self.index_prefix = index_prefix
        self.scraper = IOCScraper(
            tld_file, suppress_types=suppress_types, js_context=js_context,
        )
        self.extracted_iocs: List[ExtractedIOC] = []

    def extract(self) -> List[ExtractedIOC]:
        """
        Extract IOCs from strings and decompiled functions in analysis_results.

        Returns:
            List of ExtractedIOC objects
        """
        self.log.debug(inspect.currentframe().f_code.co_name)
        self.extracted_iocs = []

        # Extract from strings
        strings_count = self._extract_from_strings()

        # Extract from decompiled functions
        functions_count = self._extract_from_decompiled()

        # Extract from text-based artefact surfaces (JS, PowerShell, etc.)
        text_count = self._extract_from_text()

        self.log.info(
            f"Extracted {len(self.extracted_iocs)} IOCs for {self.sha256[:16]}... "
            f"(strings: {strings_count}, functions: {functions_count}, "
            f"text: {text_count})"
        )

        return self.extracted_iocs

    def _extract_from_strings(self) -> int:
        """Extract IOCs from sample's strings."""
        count = 0
        strings = self.analysis_results.get("strings", [])

        for s in strings:
            string_value = s.get("string", "")
            string_offset = s.get("string_offset", 0)

            if isinstance(string_value, bytes):
                string_value = string_value.decode('utf-8', errors='replace')

            for ioc in self.scraper.scrape(string_value, SourceType.STRING, str(string_offset)):
                self.extracted_iocs.append(ioc)
                count += 1

        return count

    def _extract_from_decompiled(self) -> int:
        """Extract IOCs from sample's decompiled functions.

        Supports both Binja format (key: "decompiled", fields: "decompiled_function",
        "decompiled_function_hash", "function_type") and APK format (key:
        "decompiled_content", fields: "decompiled_method", "decompiled_method_hash",
        "method_type").
        """
        count = 0

        # Binja format
        decompiled = self.analysis_results.get("decompiled", [])
        for func in decompiled:
            func_type = func.get("function_type", "UNKNOWN")
            if func_type in ("LIBRARY", "THUNK"):
                continue

            func_content = func.get("decompiled_function", "")
            func_hash = func.get("decompiled_function_hash", "unknown")

            if isinstance(func_content, bytes):
                func_content = func_content.decode('utf-8', errors='replace')

            for ioc in self.scraper.scrape(func_content, SourceType.DECOMPILED_FUNCTION, func_hash):
                self.extracted_iocs.append(ioc)
                count += 1

        # APK format (decompiled_content with method-level fields)
        decompiled_content = self.analysis_results.get("decompiled_content", [])
        for func in decompiled_content:
            func_type = func.get("method_type", "UNKNOWN")
            if func_type in ("LIBRARY", "THUNK"):
                continue

            func_content = func.get("decompiled_method", "")
            func_hash = func.get("decompiled_method_hash", "unknown")

            if isinstance(func_content, bytes):
                func_content = func_content.decode('utf-8', errors='replace')

            for ioc in self.scraper.scrape(func_content, SourceType.DECOMPILED_FUNCTION, func_hash):
                self.extracted_iocs.append(ioc)
                count += 1

        return count

    def _extract_from_text(self) -> int:
        """Extract IOCs from text-based artefact surfaces.

        Walks `analysis_results["text_raw"]` and `analysis_results["text_normalized"]`,
        each a list of `{"content": str, "content_hash": str}` dicts. Each
        list is routed through its own SourceType (`TEXT_RAW` /
        `TEXT_NORMALIZED`) so analysts can distinguish IOCs that were already
        present in the raw source from those exposed only after normalisation
        (deobfuscation/beautification). Generic across text-based formats —
        used by JS today, intended for PowerShell, Python, email body,
        extracted PDF/Office text in the future.
        """
        count = 0

        for key, source_type in (
            ("text_raw", SourceType.TEXT_RAW),
            ("text_normalized", SourceType.TEXT_NORMALIZED),
        ):
            for entry in self.analysis_results.get(key, []):
                content = entry.get("content", "")
                content_hash = entry.get("content_hash", "unknown")

                if isinstance(content, bytes):
                    content = content.decode('utf-8', errors='replace')

                for ioc in self.scraper.scrape(content, source_type, content_hash):
                    self.extracted_iocs.append(ioc)
                    count += 1

        return count

    def prepare_export_data(self, exporter_type: str) -> Any:
        """
        Prepare data for specific export type.

        Returns tuple for ClickHouse or list of dicts for Print/Elasticsearch.
        """
        self.log.debug(inspect.currentframe().f_code.co_name)

        if not self.extracted_iocs:
            return None

        now = datetime.now(timezone.utc)

        if exporter_type == "ClickHouseExporter":
            data = [
                [
                    self.sha256,
                    ioc.ioc_type.value,
                    ioc.ioc_value,
                    ioc.source_type.value,
                    ioc.source_identifier,
                    now,
                ]
                for ioc in self.extracted_iocs
            ]

            column_names = [
                "sha256",
                "ioc_type",
                "ioc_value",
                "source_type",
                "source_identifier",
                "extracted_at",
            ]

            column_type_names = [
                "FixedString(64)",
                "Enum8('ipv4'=1, 'ipv6'=2, 'fqdn'=3, 'url'=4, 'email'=5, 'server'=6, "
                "'hash_md5'=10, 'hash_sha1'=11, 'hash_sha256'=12, 'cve'=20, 'cwe'=21, 'cpe'=22, "
                "'crypto_btc'=30, 'crypto_eth'=31, 'crypto_xrp'=32, 'crypto_bch'=33, "
                "'crypto_ada'=34, 'crypto_substrate'=35, 'path_linux'=40, 'path_windows'=41, "
                "'registry_key'=42, 'onion'=50)",
                "String",
                "Enum8('decompiled_function'=1, 'disassembled_function'=2, 'string'=3, "
                "'text_raw'=4, 'text_normalized'=5)",
                "String",
                "DateTime64(3, 'UTC')",
            ]

            return (data, column_names, column_type_names)

        else:
            # For PrintExporter and others - return list of dicts
            return [
                {
                    "sha256": self.sha256,
                    "ioc_type": ioc.ioc_type.value,
                    "ioc_value": ioc.ioc_value,
                    "source_type": ioc.source_type.value,
                    "source_identifier": ioc.source_identifier,
                    "extracted_at": now.isoformat(),
                }
                for ioc in self.extracted_iocs
            ]

    def get_clickhouse_table(self) -> str:
        """Return the ClickHouse table name for IOCs."""
        return "redb_iocs"

    def tag(self) -> str:
        """Return the tag for this extractor."""
        return Tag.IOC.value if hasattr(Tag, 'IOC') else "ioc"

    def export_data(self) -> bool:
        """
        Export extracted IOCs to all configured exporters.

        Returns:
            True if export succeeded, False if failed, None if no data
        """
        self.log.debug(inspect.currentframe().f_code.co_name)

        # First extract the IOCs
        extracted = self.extract()

        if not extracted:
            self.log.debug("No IOCs extracted, skipping export")
            return None

        success = True

        from redb.extractors.database_exporters import PrintExporter, ClickHouseExporter

        for exporter in self.exporters:
            try:
                if isinstance(exporter, PrintExporter):
                    # For PrintExporter, pass the list of dicts
                    export_data = self.prepare_export_data("PrintExporter")
                    success &= exporter.export(export_data)

                elif isinstance(exporter, ClickHouseExporter):
                    # For ClickHouse, pass tuple with table info
                    export_data = self.prepare_export_data("ClickHouseExporter")
                    if export_data:
                        success &= exporter.export(
                            export_data,
                            table=self.get_clickhouse_table(),
                            column_names=export_data[1],
                            column_type_names=export_data[2]
                        )

            except Exception as e:
                self.log.error(f"Error exporting IOCs to {exporter.__class__.__name__}: {e}")
                success = False

        return success