H. Timothy Bunnell

45 papers A* 1A 22Misc 1Journal 8Unranked 13
YearRankTypeTitle / Venue / Authors
2026 J jnl
J. Am. Medical Informatics Assoc.
Lorna E. Thorpe, Jasmin Divers, Annemarie Hirsch, Brian S. Schwartz, Bo Cai, Jihad S. Obeid, Angela Liese, Tessa Crume, Anna Bellatorre, Jiang Bian, Yi Guo, Sarah Bost, Tianchen Lyu, Matthew T. Mefford, Matt Zhou, Eva Lustigova, Levon H. Utidjian, Mitchell Maltenfort, Patrick Hanley, Meda E. Pavkov, Marc B. Rosenman, Andrea R. Titus, Charles Bailey, Christopher B. Forrest, Levon Utidjian, Mitch Maltenfort, Amy Shah, Eneida A. Mendonça, G. Todd Alonso, Sara J. Deakyne Davies, H. Timothy Bunnell, Anne Kazak, Melody Kitzmiller, Manmohan Kamboj, Dimitri A. Christakis, Daksha Ranade, Annemarie G. Hirsch, Joseph J. Dewalle, H. Lester Kirchner, Meredith Lewis, Dione G. Mercer, Cara M. Nordberg, Amy Poissant, Brian E. Dixon, Shaun J. Grannis, Katie S. Allen, Anna Roberts, Nimish Valvi, Jeff Warvel, Ashley Wiensch, Tamara S. Hannon, Kristi Reynolds, John Chang, Don McCarthy, Rong Wei, Hui Zhou, Marc Rosenman, Lu Zhang, George Lales, Anthony Wong, Allison Zelinski, Yuan Luo, Mark Weiner, Pedro Rivera, Thomas Carton, Elizabeth Nauman, Harold P. Lehmann, Meredith Akerman, Rebecca Anthopolos, Stefanie Bendik, Sarah Conderino, Andrew Fair, Jessica Guillaume, Shahidul Islam, Alan Jacobson, David C. Lee, Chinyere Okpara, Anand Rajan, Andrea Titus, Dana Dabelea, Theresa Anderson, Rebecca Conway, Toan Ong, Jack Pattee, Shawna Burgett, Hui Shao, Elizabeth A. Shenkman, Sarah J. Bost, William T. Donahoo, William R. Hogan, Piaopiao Li, Mattia Prosperi, Yonghui Wu, Angela D. Liese, Lisa Knight, Caroline Rudisill, Jessica Stucker, Deborah Bowlby, Elaine Apperson, Alex Ewing, Giuseppina Imperatore, Deborah Rolka, Hui Xie, Ibrahim Zaganjor
2025 J jnl
CoRR
Zilong Bai, Zihan Xu, Cong Sun, Chengxi Zang, H. Timothy Bunnell, Catherine Sinfield, Jacqueline R. M. A. Maasch, Aaron Thomas Martinez, L. Charles Bailey, Mark G. Weiner, Thomas R. Campion Jr., Thomas Carton, Christopher B. Forrest, Rainu Kaushal, Fei Wang, Yifan Peng
2024 conf
ML4H@NeurIPS
Hamed Fayyaz, Mehak Gupta, Alejandra Perez Ramirez, Claudine Jurkovitz, H. Timothy Bunnell, Thao-Ly T. Phan, Rahmatollah Beheshti
2024 J jnl
CoRR
Hamed Fayyaz, Mehak Gupta, Alejandra Perez Ramirez, Claudine Jurkovitz, H. Timothy Bunnell, Thao-Ly T. Phan, Rahmatollah Beheshti
2022 A* conf
AAAI
Mehak Gupta, Raphael Poulain, Thao-Ly T. Phan, H. Timothy Bunnell, Rahmatollah Beheshti
2022 J jnl
ACM Trans. Comput. Heal.
Mehak Gupta, Thao-Ly T. Phan, H. Timothy Bunnell, Rahmatollah Beheshti
2022 conf
ML4H@NeurIPS
Hamed Fayyaz, Thao-Ly T. Phan, H. Timothy Bunnell, Rahmatollah Beheshti
2022 J jnl
CoRR
Hamed Fayyaz, Thao-Ly T. Phan, H. Timothy Bunnell, Rahmatollah Beheshti
2021 conf
BCB
Mehak Gupta, Thao-Ly T. Phan, H. Timothy Bunnell, Rahmatollah Beheshti
2021 A conf
Interspeech
Jason Lilley, H. Timothy Bunnell
2019 J jnl
CoRR
Mehak Gupta, Thao-Ly T. Phan, George Datto, H. Timothy Bunnell, Rahmatollah Beheshti
2018 A conf
INTERSPEECH
Jason Lilley, Erin L. Crowgey, H. Timothy Bunnell
2017 J jnl
J. Biomed. Informatics
Holly Antal, H. Timothy Bunnell, Suzanne M. McCahan, Christopher A. Pennington, Tim Wysocki, Kathryn V. Blake
2017 A conf
INTERSPEECH
Jason Lilley, Madhavi Vedula Ratnagiri, H. Timothy Bunnell
2017 A conf
INTERSPEECH
H. Timothy Bunnell, Jason Lilley, Kathleen McGrath
2014 A conf
INTERSPEECH
Jason Lilley, James J. Mahshie, H. Timothy Bunnell
2014 A conf
INTERSPEECH
Jason Lilley, Susan Nittrouer, H. Timothy Bunnell
2013 A conf
INTERSPEECH
Prasanna Kumar Muthukumar, Alan W. Black, H. Timothy Bunnell
2013 A conf
INTERSPEECH
Shanqing Cai, H. Timothy Bunnell, Rupal Patel
2012 Misc conf
ICASSP
Alan W. Black, H. Timothy Bunnell, Ying Dou, Prasanna Kumar Muthukumar, Florian Metze, Daniel Perry, Tim Polzehl, Kishore Prahallad, Stefan Steidl, Callie Vaughn
2012 A conf
INTERSPEECH
Christian DiCanio, Hosung Nam, Douglas H. Whalen, H. Timothy Bunnell, Jonathan D. Amith, Rey Castillo García
2012 J jnl
J. Phonetics
Laura Spinu, Irene Vogel, H. Timothy Bunnell
2012 A conf
INTERSPEECH
Kyoko Nagao, Mark Paullin, James B. Polikoff, Jason Lilley, H. Timothy Bunnell
2012 A conf
INTERSPEECH
Kyoko Nagao, Mark Paullin, Vilena Livinsky, James B. Polikoff, Linda D. Vallino, Thierry G. Morlet, N. Carolyn Schanen, H. Timothy Bunnell
2012 A conf
INTERSPEECH
Ann K. Syrdal, H. Timothy Bunnell, Susan R. Hertz, Taniya Mishra, Murray F. Spiegel, Corine A. Bickley, Deborah Rekart, Matthew J. Makashay
2011 A conf
INTERSPEECH
H. Timothy Bunnell, Jason Lilley, Sigfrid D. Soli, Ivan Pal
2010 conf
SSW
H. Timothy Bunnell
2010 conf
Blizzard Challenge
H. Timothy Bunnell, Jason Lilley, Christopher A. Pennington, Bill Moyers, James B. Polikoff
2009 A conf
INTERSPEECH
Irene Vogel, Arild Hestvik, H. Timothy Bunnell, Laura Spinu
2009 A conf
ASSETS
Camil Jreige, Rupal Patel, H. Timothy Bunnell
2008 conf
ACL (Demo Papers)
Debra Yarrington, John Gray, Christopher A. Pennington, H. Timothy Bunnell, Allegra Cornaglia, Jason Lilley, Kyoko Nagao, James B. Polikoff
2008 A conf
INTERSPEECH
H. Timothy Bunnell, Jason Lilley
2007 conf
SSW
H. Timothy Bunnell, Jason Lilley
2007 A conf
INTERSPEECH
H. Timothy Bunnell, N. Carolyn Schanen, Linda D. Vallino, Thierry G. Morlet, James B. Polikoff, Jennette D. Driscoll, James T. Mantell
2006 A conf
INTERSPEECH
H. Timothy Bunnell, James B. Polikoff
2005 A conf
ASSETS
Debra Yarrington, Christopher A. Pennington, John Gray, H. Timothy Bunnell
2005 A conf
INTERSPEECH
H. Timothy Bunnell, Christopher A. Pennington, Debra Yarrington, John Gray
2004 A conf
INTERSPEECH
H. Timothy Bunnell, James B. Polikoff, Jane McNicholas
2000 A conf
INTERSPEECH
H. Timothy Bunnell, Debra Yarrington, James B. Polikoff
1998 conf
SSW
H. Timothy Bunnell, Steven R. Hoskins, Debra Yarrington
1998 conf
ICSLP
H. Timothy Bunnell, Steven R. Hoskins, Debra Yarrington
1997 conf
EUROSPEECH
Xavier Menéndez-Pidal, James B. Polikoff, H. Timothy Bunnell
1996 conf
ICSLP
Xavier Menéndez-Pidal, James B. Polikoff, Shirley M. Peters, Jennie E. Leonzio, H. Timothy Bunnell
1995 conf
EUROSPEECH
Debra Yarrington, H. Timothy Bunnell, Gene Ball
1994 conf
SSW
H. Timothy Bunnell, Debra Yarrington, Kenneth E. Barner
redb/extractors/extractor.py
← Index redb/extractors/extractor.py python
import hashlib
import inspect
from abc import ABCMeta, abstractmethod
from dataclasses import asdict
from functools import cached_property
from datetime import datetime, timezone
import math
from typing import Counter, List, Optional, Dict, Any, Tuple
from redb import settings
from redb.models.dataclasses import Hash
from .database_exporters import DatabaseExporter, ElasticsearchExporter, ClickHouseExporter, PrintExporter

class Extractor(metaclass=ABCMeta):
    def __init__(
        self,
        filepath: str,
        log: Any,
        exporters: Optional[List[DatabaseExporter]] = None,
        index_prefix: Optional[str] = None,
        source: Optional[str] = None,
        elastic_index: Optional[str] = None,
        known_benign: bool = False,
        known_malicious: bool = False,
        precomputed_hashes: Optional[Dict[str, str]] = None,
    ):
        self.log = log
        self.log.debug(f"Creating {self.__class__.__name__}")
        self.filepath = filepath
        self.source = source
        self.exporters = exporters or []
        self.index_prefix = index_prefix if index_prefix else settings.ELASTIC_BINARIES_COLLECTION
        self.elastic_index = self.index_prefix + (elastic_index if elastic_index else "")
        self.known_benign = known_benign
        self.known_malicious = known_malicious

        # Use precomputed hashes if provided (e.g., from machofile, pefile)
        # Otherwise compute them from binary
        if precomputed_hashes:
            self.md5 = precomputed_hashes.get('md5') or precomputed_hashes.get('MD5')
            self.sha1 = precomputed_hashes.get('sha1') or precomputed_hashes.get('SHA1')
            self.sha256 = precomputed_hashes.get('sha256') or precomputed_hashes.get('SHA256')
        else:
            self.md5 = hashlib.md5(self.binary).hexdigest()
            self.sha1 = hashlib.sha1(self.binary).hexdigest()
            self.sha256 = hashlib.sha256(self.binary).hexdigest()
        self.hash = Hash(self.md5, self.sha1, self.sha256)

    @cached_property
    def binary(self):
        with open(self.filepath, "rb") as f:
            data = f.read()
        return data

    @property
    @abstractmethod
    def tag(self):
        pass

    @abstractmethod
    def extract(self):
        """
        this method defines the extracted data
        """

    @staticmethod
    def process_binary_string(s):
        # Remove \x00 padding
        s = s.rstrip(b"\x00")

        # Check if there are any non-printable characters
        has_non_printable = any(byte < 32 or byte > 126 for byte in s)

        if not has_non_printable:
            # If all characters are printable, decode the string
            return s.decode()
        else:
            # If there are non-printable characters, represent them as \xDD
            return "".join(
                [
                    f"\\x{byte:02x}" if byte < 32 or byte > 126 else chr(byte)
                    for byte in s
                ]
            )

    @staticmethod
    def remove_non_utf8(binary_string):
        decoded = b""
        for i in range(len(binary_string)):
            try:
                # Try to decode each byte
                char = binary_string[i : i + 1].decode("utf-8")
                decoded += char.encode("utf-8")
            except UnicodeDecodeError:
                # Skip this byte if it can't be decoded
                continue
        return decoded

    def calculate_entropy(self, data):
        """Calculate the entropy of a chunk of data.
        Based on pefile.SectionStructure.entropy_H.
        """
        # self.log.debug(inspect.currentframe().f_code.co_name)
        if not data:
            return 0.0

        if type(data) == str:
            counts = Counter(data)
            frequencies = ((i / len(data)) for i in counts.values())
            return - sum(f * math.log(f, 2) for f in frequencies)
        else:
            occurences = Counter(bytearray(data))
            entropy = 0
            for x in occurences.values():
                p_x = float(x) / len(data)
                entropy -= p_x * math.log(p_x, 2)
            return entropy

    @abstractmethod
    def prepare_export_data(self, exporter_type: str) -> Tuple[List[Any], List[str], List[str]]:
        """
        Prepare data for specific export type
        Returns:
            Tuple containing:
            - data: List of values to insert
            - column_names: List of column names
            - column_type_names: List of column types
        """
        pass

    def export_data(self):
        """Export data to all configured exporters

        Returns:
            True: Export succeeded
            False: Export failed (actual error)
            None: No data to export (not an error, e.g., no overlay, no signature)
        """
        self.log.debug(inspect.currentframe().f_code.co_name)
        success = True
        extracted_data = self.extract()

        if extracted_data is None:
            self.log.debug("extract() returned None, skipping export")
            return None  # No data to export, not a failure
            
        for exporter in self.exporters:
            if isinstance(exporter, PrintExporter):
                # For PrintExporter, we pass the extracted data directly
                success &= exporter.export(extracted_data)
            else:
                # Get the data prepared for this specific exporter type
                export_data = self.prepare_export_data(exporter.__class__.__name__)
                
                if export_data is None:
                    self.log.debug(f"prepare_export_data returned None for {exporter.__class__.__name__}")
                    return False
                
                if isinstance(exporter, ElasticsearchExporter):
                    success &= exporter.export(
                        export_data,
                        index=self.elastic_index,
                        tag=self.tag(),
                        hashes=asdict(self.hash),
                        known_benign=self.known_benign,
                        known_malicious=self.known_malicious
                    )
                elif isinstance(exporter, ClickHouseExporter):
                    # For ClickHouse, we need to pass the table name and the prepared data
                    success &= exporter.export(
                        export_data,
                        table=self.get_clickhouse_table(),
                        # Add these parameters explicitly
                        column_names=export_data[1] if isinstance(export_data, tuple) else None,
                        column_type_names=export_data[2] if isinstance(export_data, tuple) else None
                    )
                
        return success

    @abstractmethod
    def get_clickhouse_table(self) -> str:
        """Return the appropriate ClickHouse table name"""
        pass
    # def export_to_elastic(self, list_of_dataclasses, tag=None):
    #     self.log.debug(inspect.currentframe().f_code.co_name)

    #     if not isinstance(list_of_dataclasses, list):
    #         self.log.error("Called export_to_elastic wrongly")

    #     now_t = datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M:%S")
    #     for dataclass_ in list_of_dataclasses:
    #         if not settings.ELASTIC_CLIENT.ping():
    #             self.log.error("[CONNECTION ERROR] ping to elastic failed")
            
    #         # Convert dataclass to dict and filter out None values
    #         document = {k: v for k, v in asdict(dataclass_).items() if v is not None}
            
    #         if tag:
    #             document["tag"] = [tag, self.tag()]
    #         else:
    #             document["tag"] = self.tag()
    #         hashes = asdict(self.hash)
    #         document |= hashes
    #         document["timestamp_utc"] = now_t
    #         # document["source"] = self.source
    #         document["known_benign"] = self.known_benign
    #         document["known_malicious"] = self.known_malicious

    #         if "_id" in document:
    #             tmp_id = document.pop("_id") + document["sha256"]
    #             _id = hashlib.sha256(tmp_id.encode()).hexdigest()
    #         else:
    #             _id = document["sha256"]

    #         # self.log.debug(f"[DEBUG] about to export {type(document)} {document}")
    #         try:
    #             doc_dump = json.dumps(document)
    #         except TypeError as e:
    #             self.log.error(
    #                 f"Failed export of document. " f"full document: {document}"
    #             )
    #             raise e

    #         # body={"doc": doc_dump,
    #         #       "doc_as_upsert": True  # Create the document if it doesn't exist
    #         # }

    #         # Check if the index exists, and create it if it doesn't
    #         # if not settings.ELASTIC_CLIENT.indices.exists(index=self.elastic_index):
    #         #     settings.ELASTIC_CLIENT.indices.create(index=self.elastic_index)
    #         # pprint(doc_dump) #DEBUG
    #         # print("[DEBUG] _id: " + _id)
    #         # print("[DEBUG] index: " + self.elastic_index)
    #         settings.ELASTIC_CLIENT.index(
    #             index=self.elastic_index, id=_id, document=doc_dump
    #         )