J. Harry Caufield

26 papers A* 2Journal 21Unranked 3
YearRankTypeTitle / Venue / Authors
2025 J jnl
Database J. Biol. Databases Curation
Harshad Hegde, Jennifer Vendetti, Damien Goutte-Gattat, J. Harry Caufield, John B. Graybeal, Nomi L. Harris, Naouel Karam, Christian Kindermann, Nicolas Matentzoglu, James A. Overton, Mark A. Musen, Christopher J. Mungall
2025 conf
TPDL (Short Papers and Workshops)
Tim Alamenciak, Carlos Alberto Arnillas, J. Harry Caufield, Katherine Compton, Kian Drew, Robert Frühstückl, Tina Heger, Birgitta König-Ries, Chris Mungall, Sierra A. T. Moxon, Justin T. Reese, Jordan Tardif, Lars Vogt
2025 J jnl
CoRR
Daniel R. Korn, Patrick Golden, Aaron Odell, Katherina G. Cortes, Shilpa Sundar, Kevin Schaper, Sarah Gehrke, Corey Cox, J. Harry Caufield, Justin T. Reese, Evan Morris, Christopher J. Mungall, Melissa A. Haendel
2025 J jnl
CoRR
Sierra A. T. Moxon, Harold Solbrig, Nomi L. Harris, Patrick Kalita, Mark A. Miller, Sujay Patil, Kevin Schaper, Chris Bizon, J. Harry Caufield, Silvano Cirujano Cuesta, Corey Cox, Frank Dekervel, Damion M. Dooley, William D. Duncan, Tim Fliss, Sarah Gehrke, Adam S. L. Graefe, Harshad Hegde, AJ Ireland, Julius O. B. Jacobsen, Madan Krishnamurthy, Carlo Kroll, David Linke, Ryan Ly, Nicolas Matentzoglu, James A. Overton, Jonny L. Saunders, Deepak R. Unni, Gaurav Vaidya, Wouter-Michiel A. M. Vierdag, LinkML Community Contributors, Oliver Ruebel, Christopher G. Chute, Matthew H. Brush, Melissa A. Haendel, Christopher J. Mungall
2025 J jnl
CoRR
J. Harry Caufield, Satrajit Ghosh, Sek Wong Kong, Jillian Parker, Nathan C. Sheffield, Bhavesh Patel, Andrew Williams, Timothy Clark, Monica C. Munoz-Torres
2024 J jnl
CoRR
Harshad Hegde, Jennifer Vendetti, Damien Goutte-Gattat, J. Harry Caufield, John B. Graybeal, Nomi L. Harris, Naouel Karam, Christian Kindermann, Nicolas Matentzoglu, James A. Overton, Mark A. Musen, Christopher J. Mungall
2024 J jnl
BMC Medical Informatics Decis. Mak.
Tudor Groza, J. Harry Caufield, Dylan Gration, Gareth Baynam, Melissa A. Haendel, Peter N. Robinson, Christopher J. Mungall, Justin T. Reese
2024 J jnl
CoRR
J. Harry Caufield, Carlo Kroll, Shawn T. O'Neil, Justin T. Reese, Marcin P. Joachimiak, Harshad Hegde, Nomi L. Harris, Madan Krishnamurthy, James A. McLaughlin, Damian Smedley, Melissa A. Haendel, Peter N. Robinson, Christopher J. Mungall
2024 conf
SEBD
Emanuele Cavalleri, Mauricio Soto Gomez, Ali Pashaeibarough, Dario Malchiodi, J. Harry Caufield, Justin T. Reese, Christopher J. Mungall, Peter N. Robinson, Elena Casiraghi, Giorgio Valentini, Marco Mesiti
2024 conf
VLDB Workshops
Emanuele Cavalleri, Mauricio Soto Gomez, Ali Pashaeibarough, Dario Malchiodi, J. Harry Caufield, Justin T. Reese, Chris Mungall, Peter N. Robinson, Elena Casiraghi, Giorgio Valentini, Marco Mesiti
2024 J jnl
Bioinform.
J. Harry Caufield, Harshad Hegde, Vincent Emonet, Nomi L. Harris, Marcin P. Joachimiak, Nicolas Matentzoglu, HyeongSik Kim, Sierra A. T. Moxon, Justin T. Reese, Melissa A. Haendel, Peter N. Robinson, Christopher J. Mungall
2024 J jnl
Appl. Ontology
Marcin P. Joachimiak, Mark A. Miller, J. Harry Caufield, Ryan Ly, Nomi L. Harris, Andrew J. Tritt, Christopher J. Mungall, Kristofer E. Bouchard
2024 J jnl
CoRR
Marcin P. Joachimiak, Mark A. Miller, J. Harry Caufield, Ryan Ly, Nomi L. Harris, Andrew J. Tritt, Christopher J. Mungall, Kristofer E. Bouchard
2023 J jnl
CoRR
Tudor Groza, J. Harry Caufield, Dylan Gration, Gareth Baynam, Melissa A. Haendel, Peter N. Robinson, Chris Mungall, Justin T. Reese
2023 J jnl
CoRR
Marcin P. Joachimiak, J. Harry Caufield, Nomi L. Harris, HyeongSik Kim, Christopher J. Mungall
2023 J jnl
CoRR
J. Harry Caufield, Tim E. Putman, Kevin Schaper, Deepak R. Unni, Harshad Hegde, Tiffany J. Callahan, Luca Cappelletti, Sierra A. T. Moxon, Vida Ravanmehr, Seth Carbon, Lauren E. Chan, Katherina G. Cortes, Kent A. Shefchek, Glass Elsarboukh, James P. Balhoff, Tommaso Fontana, Nicolas Matentzoglu, Richard M. Bruskiewich, Anne E. Thessen, Nomi L. Harris, Monica C. Munoz-Torres, Melissa A. Haendel, Peter N. Robinson, Marcin P. Joachimiak, Christopher J. Mungall, Justin T. Reese
2023 J jnl
Bioinform.
J. Harry Caufield, Tim E. Putman, Kevin Schaper, Deepak R. Unni, Harshad Hegde, Tiffany J. Callahan, Luca Cappelletti, Sierra A. T. Moxon, Vida Ravanmehr, Seth Carbon, Lauren E. Chan, Katherina G. Cortes, Kent A. Shefchek, Glass Elsarboukh, James P. Balhoff, Tommaso Fontana, Nicolas Matentzoglu, Richard M. Bruskiewich, Anne E. Thessen, Nomi L. Harris, Monica C. Munoz-Torres, Melissa A. Haendel, Peter N. Robinson, Marcin P. Joachimiak, Christopher J. Mungall, Justin T. Reese
2023 J jnl
CoRR
Nicolas Matentzoglu, J. Harry Caufield, Harshad B. Hegde, Justin T. Reese, Sierra A. T. Moxon, HyeongSik Kim, Nomi L. Harris, Melissa A. Haendel, Christopher J. Mungall
2023 J jnl
CoRR
J. Harry Caufield, Harshad Hegde, Vincent Emonet, Nomi L. Harris, Marcin P. Joachimiak, Nicolas Matentzoglu, HyeongSik Kim, Sierra A. T. Moxon, Justin T. Reese, Melissa A. Haendel, Peter N. Robinson, Christopher J. Mungall
2021 A* conf
ICDE
Yichao Zhou, Wei-Ting Chen, Bowen Zhang, David Lee, J. Harry Caufield, Kai-Wei Chang, Yizhou Sun, Peipei Ping, Wei Wang
2021 J jnl
CoRR
Yichao Zhou, Wei-Ting Chen, Bowen Zhang, David Lee, J. Harry Caufield, Kai-Wei Chang, Yizhou Sun, Peipei Ping, Wei Wang
2021 J jnl
CoRR
Yichao Zhou, Chelsea Ju, J. Harry Caufield, Kevin Shih, Calvin Yu-Chian Chen, Yizhou Sun, Kai-Wei Chang, Peipei Ping, Wei Wang
2021 A* conf
AAAI
Yichao Zhou, Yu Yan, Rujun Han, J. Harry Caufield, Kai-Wei Chang, Yizhou Sun, Peipei Ping, Wei Wang
2020 J jnl
CoRR
Yichao Zhou, Yu Yan, Rujun Han, J. Harry Caufield, Kai-Wei Chang, Yizhou Sun, Peipei Ping, Wei Wang
2017 J jnl
BMC Bioinform.
J. Harry Caufield, Christopher Wimble, Shary Semarjit, Stefan Wuchty, Peter Uetz
2015 J jnl
PLoS Comput. Biol.
J. Harry Caufield, Marco Abreu, Christopher Wimble, Peter Uetz
tests/scripts/test_macho_extractors_local.py
← Index tests/scripts/test_macho_extractors_local.py python
#!/usr/bin/env python3
"""
Local MachO extractors test - test actual extractors without server connections
"""

import sys
import os
import logging
import hashlib
from datetime import datetime, timezone
from typing import Dict, Any, Optional
from unittest.mock import Mock, patch

# Add the redb directory to the path so we can import the extractors
sys.path.insert(0, os.path.join(os.path.dirname(__file__), 'redb'))

# Mock settings to avoid database connections
with patch.dict('os.environ', {'REDB_ENV': 'test'}):
    # Import the actual extractors
    from redb.extractors.macho_extractors.macho_features import MachOFeaturesExtractor
    from redb.extractors.macho_extractors.macho_segments import MachOSegmentExtractor
    from redb.extractors.macho_extractors.macho_imports import MachOImportExtractor
    from redb.extractors.macho_extractors.macho_exports import MachOExportExtractor
    from redb.extractors.macho_extractors.macho_dylibs import MachODylibExtractor
    from redb.extractors.macho_extractors.macho_signature import MachOSignatureExtractor
    from redb.extractors.macho_extractors.macho_universal import MachOUniversalExtractor

def setup_logging():
    """Setup basic logging configuration."""
    logging.basicConfig(
        level=logging.INFO,
        format='%(asctime)s - %(levelname)s - %(message)s'
    )
    return logging.getLogger(__name__)

class MockLogger:
    """Mock logger for testing without RedB dependencies."""
    def __init__(self):
        self.logger = logging.getLogger(__name__)
    
    def debug(self, msg):
        self.logger.debug(msg)
    
    def info(self, msg):
        self.logger.info(msg)
    
    def warning(self, msg):
        self.logger.warning(msg)
    
    def error(self, msg):
        self.logger.error(msg)

class MockExporter:
    """Mock exporter that does nothing."""
    def export(self, data):
        pass

def create_mock_extractor(extractor_class, filepath):
    """Create an extractor instance with mocked dependencies."""
    log = MockLogger()
    
    # Mock the exporters to avoid database connections
    mock_exporters = [MockExporter()]
    
    # Create the extractor with mocked dependencies
    extractor = extractor_class(
        filepath=filepath,
        log=log,
        exporters=mock_exporters,
        index_prefix="test",
        elastic_index="test_macho",
        known_benign=False,
        known_malicious=False
    )
    
    return extractor

def test_extractor(extractor_class, filepath, extractor_name):
    """Test a specific extractor."""
    log = setup_logging()
    log.info(f"Testing {extractor_name}...")
    
    try:
        # Create the extractor with mocked dependencies
        extractor = create_mock_extractor(extractor_class, filepath)
        
        if not extractor.macho:
            log.error(f"{extractor_name}: No MachO object created")
            return False

        # Test the actual extract method
        result = extractor.extract()
        
        if result:
            log.info(f"{extractor_name}: Successfully extracted data")
            print(f"\n{'='*60}")
            print(f"=== {extractor_name.upper()} RESULTS ===")
            print(f"{'='*60}")
            
            if isinstance(result, list):
                print(f"📊 Extracted {len(result)} items")
                print()
                for i, item in enumerate(result):
                    print(f"📦 Item {i+1}:")
                    print(f"   {'─'*40}")
                    # Handle dataclass objects in lists
                    if hasattr(item, '__dataclass_fields__'):
                        from dataclasses import asdict
                        item_dict = asdict(item)
                        for key, value in item_dict.items():
                            if isinstance(value, (list, dict)):
                                if isinstance(value, list):
                                    print(f"   🔹 {key}: List with {len(value)} items")
                                    print(f"      {value}")
                                elif isinstance(value, dict):
                                    print(f"   🔹 {key}: Dict with {len(value)} keys")
                                    for k, v in value.items():
                                        print(f"      {k}: {v}")
                            else:
                                print(f"   🔹 {key}: {value}")
                    else:
                        # Regular dict
                        for key, value in item.items():
                            if isinstance(value, (list, dict)):
                                if isinstance(value, list):
                                    print(f"   🔹 {key}: List with {len(value)} items")
                                    print(f"      {value}")
                                elif isinstance(value, dict):
                                    print(f"   🔹 {key}: Dict with {len(value)} keys")
                                    for k, v in value.items():
                                        print(f"      {k}: {v}")
                            else:
                                print(f"   🔹 {key}: {value}")
                    print()
            else:
                print("📊 Extracted data:")
                print()
                # Handle dataclass objects
                if hasattr(result, '__dataclass_fields__'):
                    # It's a dataclass, use dataclasses.asdict
                    from dataclasses import asdict
                    result_dict = asdict(result)
                    for key, value in result_dict.items():
                        if isinstance(value, (list, dict)):
                            if isinstance(value, list):
                                print(f"🔹 {key}: List with {len(value)} items")
                                print(f"   {value}")
                            elif isinstance(value, dict):
                                print(f"🔹 {key}: Dict with {len(value)} keys")
                                for k, v in value.items():
                                    print(f"   {k}: {v}")
                        else:
                            print(f"🔹 {key}: {value}")
                        print()
                else:
                    # It's a regular dict
                    for key, value in result.items():
                        if isinstance(value, (list, dict)):
                            if isinstance(value, list):
                                print(f"🔹 {key}: List with {len(value)} items")
                                print(f"   {value}")
                            elif isinstance(value, dict):
                                print(f"🔹 {key}: Dict with {len(value)} keys")
                                for k, v in value.items():
                                    print(f"   {k}: {v}")
                        else:
                            print(f"🔹 {key}: {value}")
                        print()
        else:
            log.warning(f"{extractor_name}: No data extracted")
            print(f"\n❌ {extractor_name}: No data extracted")
        
        return True
        
    except Exception as e:
        log.error(f"{extractor_name}: Error during extraction: {e}")
        import traceback
        traceback.print_exc()
        return False

def main():
    """Main function."""
    if len(sys.argv) < 2:
        print("Usage: python test_macho_extractors_local.py <macho_file> [extractor_name]")
        print("\nAvailable extractors:")
        print("  features    - Basic MachO header and metadata")
        print("  segments    - Segment information and analysis")
        print("  imports     - Imported functions and libraries")
        print("  exports     - Exported symbols")
        print("  dylibs      - Dynamic library dependencies")
        print("  signature   - Code signing information")
        print("  universal   - FAT/Universal binary information")
        print("  all         - Test all extractors")
        sys.exit(1)
    
    filepath = sys.argv[1]
    extractor_name = sys.argv[2] if len(sys.argv) > 2 else "all"
    
    if not os.path.exists(filepath):
        print(f"File not found: {filepath}")
        sys.exit(1)
    
    # Define extractors
    extractors = {
        'features': (MachOFeaturesExtractor, "MachO Features"),
        'segments': (MachOSegmentExtractor, "MachO Segments"),
        'imports': (MachOImportExtractor, "MachO Imports"),
        'exports': (MachOExportExtractor, "MachO Exports"),
        'dylibs': (MachODylibExtractor, "MachO Dylibs"),
        'signature': (MachOSignatureExtractor, "MachO Code Signature"),
        'universal': (MachOUniversalExtractor, "MachO Universal/FAT"),
    }
    
    if extractor_name == "all":
        print(f"Testing all extractors with file: {filepath}")
        success_count = 0
        for name, (extractor_class, display_name) in extractors.items():
            if test_extractor(extractor_class, filepath, display_name):
                success_count += 1
            print("-" * 50)
        
        print(f"\n✅ {success_count}/{len(extractors)} extractors completed successfully")
        
    elif extractor_name in extractors:
        extractor_class, display_name = extractors[extractor_name]
        success = test_extractor(extractor_class, filepath, display_name)
        
        if success:
            print(f"\n✅ {display_name} testing completed successfully")
        else:
            print(f"\n❌ {display_name} testing failed")
            sys.exit(1)
    else:
        print(f"Unknown extractor: {extractor_name}")
        print("Available extractors:", ", ".join(extractors.keys()) + ", all")
        sys.exit(1)

if __name__ == "__main__":
    main()