Veeky Baths

39 papers A* 1A 1B 4C 2Journal 20Unranked 11
YearRankTypeTitle / Venue / Authors
2025 J jnl
IEEE Access
Yesoda Bhargava, S. V. Sumanth, Anirban Deshmukh, Veeky Baths
2025 J jnl
CoRR
Niharika Tewari, Nguyen Linh Dan Le, Mujie Liu, Jing Ren, Ziqi Xu, Tabinda Sarwar, Veeky Baths, Feng Xia
2025 J jnl
IEEE Access
Kumar Paritosh, Shubhangi K. Gawali, Veeky Baths, Neena Goveas
2025 A conf
WACV
Harini S. I, Somesh Singh, Yaman Kumar Singla, Aanisha Bhattacharyya, Veeky Baths, Changyou Chen, Rajiv Ratn Shah, Balaji Krishnamurthy
2025 A* conf
ICLR
Somesh Kumar Singh, Harini S. I, Yaman Kumar Singla, Changyou Chen, Rajiv Ratn Shah, Veeky Baths, Balaji Krishnamurthy
2024 J jnl
IEEE Access
Yesoda Bhargava, Kanthi Kumar Kattupalli, Veeky Baths
2024 J jnl
CoRR
Hugo Laurençon, Yesoda Bhargava, Riddhi Zantye, Charbel-Raphaël Ségerie, Johann Lussange, Veeky Baths, Boris Gutkin
2024 J jnl
CoRR
Somesh Singh, Harini S. I, Yaman Kumar Singla, Veeky Baths, Rajiv Ratn Shah, Changyou Chen, Balaji Krishnamurthy
2024 J jnl
Frontiers Comput. Neurosci.
Sayani Mallick, Veeky Baths
2023 B conf
SMC
Ritwik Jain, Prakhar Jaiman, Veeky Baths
2023 J jnl
CoRR
Harini S. I, Somesh Singh, Yaman Kumar Singla, Aanisha Bhattacharyya, Veeky Baths, Changyou Chen, Rajiv Ratn Shah, Balaji Krishnamurthy
2023 conf
ICACDS
Merina Dhara, Veeky Baths, Aiswarya Subramanian
2023 conf
ICMLDE
Yesoda Bhargava, Sandesh Kumar Shetty, Veeky Baths
2022 conf
ICCAI
Yash Kumar, Piyush Maheshwari, Shreyansh Joshi, Veeky Baths
2022 conf
ADMA (2)
Shounak Naik, Rajaswa Patil, Swati Agarwal, Veeky Baths
2022 J jnl
CoRR
Shounak Naik, Rajaswa Patil, Swati Agarwal, Veeky Baths
2022 J jnl
Neural Networks
Ajay Subramanian, Sharad Chitlangia, Veeky Baths
2021 J jnl
Big Data Cogn. Comput.
Bhargav Prakash, Gautam Kumar Baboo, Veeky Baths
2021 conf
ISCMI
Akshay Valsaraj, Ithihas Madala, Nikhil Garg, Veeky Baths
2021 J jnl
CoRR
Akshay Valsaraj, Ithihas Madala, Nikhil Garg, Veeky Baths
2021 J jnl
CoRR
Mahak Kothari, Shreyansh Joshi, Adarsh Nandanwar, Aadetya Jaiswal, Veeky Baths
2021 J jnl
CoRR
Yash Kumar, Piyush Maheshwari, Shreyansh Joshi, Veeky Baths
2021 J jnl
CoRR
Arijit Gupta, Rajaswa Patil, Veeky Baths
2021 J jnl
CoRR
Rajaswa Patil, Jasleen Dhillon, Siddhant Mahurkar, Saumitra Kulkarni, Manav Malhotra, Veeky Baths
2021 B conf
SMC
Lizy Kanungo, Nikhil Garg, Anish Bhobe, Smit Rajguru, Veeky Baths
2021 J jnl
CoRR
Lizy Kanungo, Nikhil Garg, Anish Bhobe, Smit Rajguru, Veeky Baths
2020 B conf
SMC
Kshitij Chhabra, Pranay Mathur, Veeky Baths
2020 conf
SemEval@COLING
Rajaswa Patil, Veeky Baths
2020 J jnl
CoRR
Rajaswa Patil, Veeky Baths
2020 C conf
CW
Akshay Valsaraj, Ithihas Madala, Nikhil Garg, Mohit Patil, Veeky Baths
2020 J jnl
CoRR
Ajay Subramanian, Sharad Chitlangia, Veeky Baths
2020 conf
MediaEval
Alish Dipani, Gaurav Iyer, Veeky Baths
2019 C conf
ICSOFT
Shivam, Nilanjana Goswami, Veeky Baths, Soumyadip Bandyopadhyay
2019 conf
ICBSP
Prabhav Mehra, Rajee Gupta, Abhishek Mahajan, Veeky Baths
2019 conf
ICCI*CC
Mohit Patil, Nikhil Garg, Lizy Kanungo, Veeky Baths
2017 conf
EMBC
Advait Balaji, Aparajita Haldar, Keshav Patil, Sai Ruthvik Thandayam, C. A. Valliappan, Mayur Jartarkar, Veeky Baths
2016 conf
MobiHealth
C. A. Valliappan, Advait Balaji, Sai Ruthvik Thandayam, Piyush Dhingra, Veeky Baths
2014 J jnl
CoRR
Kratarth Goel, Raunaq Vohra, Anant Kamath, Veeky Baths
2014 B conf
SMC
Kratarth Goel, Raunaq Vohra, Anant Kamath, Veeky Baths
redb/extractors/macho_extractor.py
← Index redb/extractors/macho_extractor.py python
import logging
from abc import ABCMeta, abstractmethod
import inspect
import sys
import os

import machofile

from redb.extractors.extractor import Extractor

logger = logging.getLogger(__name__)


@abstractmethod
class MachOExtractor(Extractor, metaclass=ABCMeta):

    def __init__(
        self,
        filepath,
        log,
        exporters=None,
        index_prefix=None,
        elastic_index=None,
        known_benign=False,
        known_malicious=False,
        macho=None,
    ):
        # Read binary and parse machofile BEFORE calling super().__init__
        # This avoids reading the file twice
        with open(filepath, "rb") as f:
            binary_data = f.read()

        # Parse machofile with binary data
        self.macho = macho if macho else self._generate_machofile_object(binary_data)

        # Extract hashes from machofile to pass to parent
        precomputed_hashes = None
        if self.macho:
            try:
                general_info = self.macho.get_general_info()
                if general_info:
                    # For FAT binaries, get_general_info() returns dict with 'fat' key
                    # For single-arch, it returns the info directly
                    if 'fat' in general_info:
                        fat_info = general_info['fat']
                        precomputed_hashes = {
                            'MD5': fat_info.get('MD5'),
                            'SHA1': fat_info.get('SHA1'),
                            'SHA256': fat_info.get('SHA256'),
                        }
                    else:
                        precomputed_hashes = {
                            'MD5': general_info.get('MD5'),
                            'SHA1': general_info.get('SHA1'),
                            'SHA256': general_info.get('SHA256'),
                        }
            except Exception as e:
                logger.debug(f"Could not get hashes from machofile: {e}")

        super().__init__(
            filepath,
            log,
            exporters,
            index_prefix,
            elastic_index,
            known_benign,
            known_malicious,
            precomputed_hashes=precomputed_hashes,
        )

        # Store binary data so base class doesn't re-read
        self._binary_data = binary_data

    @property
    def binary(self):
        """Override to use already-read binary data."""
        return self._binary_data

    def _generate_machofile_object(self, binary_data):
        """Generate and parse a machofile object from binary data."""
        macho = None
        try:
            macho = machofile.UniversalMachO(data=binary_data)
            if not macho:
                raise Exception("Empty file?")

            # Parse the MachO object once during initialization
            macho.parse()

        except Exception as e:
            logger.error(f"Format error parsing MachO: {e}")
        return macho

    # def _is_macho_file(self):
    #     """Check if the file is a valid Mach-O binary."""
    #     try:
    #         if not self.macho:
    #             return False
            
    #         # For Universal/FAT binaries, check if any architecture is valid
    #         if hasattr(self.macho, 'is_fat') and self.macho.is_fat:
    #             return len(self.macho.architectures) > 0
    #         else:
    #             # Single architecture binary
    #             return hasattr(self.macho, 'macho') and self.macho.macho is not None
    #     except Exception as e:
    #         self.log.error(f"Error checking Mach-O file: {e}")
    #         return False

    def _is_signed(self):
        """Check if the Mach-O binary is code signed using new API."""
        try:
            if not self.macho:
                return False

            # Get architectures using new API
            architectures = self.macho.get_architectures()

            # For each architecture, check if signed
            for arch in architectures:
                try:
                    signature_info = self.macho.get_code_signature_info(arch=arch)
                    if signature_info and signature_info.get('signed', False):
                        return True
                except Exception:
                    continue

            return False
        except Exception as e:
            self.log.error(f"Error checking Mach-O signature: {e}")
            return False

    def _get_architectures(self):
        """Get list of architectures in the Mach-O binary using new API."""
        try:
            if not self.macho:
                return []

            # Use new API method
            architectures = self.macho.get_architectures()
            return architectures if architectures else []
        except Exception as e:
            self.log.error(f"Error getting architectures: {e}")
            return []

    # def _get_macho_for_arch(self, arch_name=None):
    #     """Get MachO instance for specific architecture or default."""
    #     try:
    #         if not self.macho:
    #             return None
            
    #         if hasattr(self.macho, 'is_fat') and self.macho.is_fat:
    #             if arch_name:
    #                 return self.macho.architectures.get(arch_name)
    #             else:
    #                 # Return first available architecture
    #                 return next(iter(self.macho.architectures.values())) if self.macho.architectures else None
    #         else:
    #             # Single architecture binary
    #             return self.macho.macho if hasattr(self.macho, 'macho') else None
    #     except Exception as e:
    #         self.log.error(f"Error getting MachO for architecture: {e}")
    #         return None

    # def _get_formatted_header_values(self, header):
    #     """Get both raw and human-readable header values."""
    #     try:
    #         macho_instance = self._get_macho_for_arch()
    #         if not macho_instance:
    #             return None
            
    #         # Parse the MachO if not already parsed
    #         if not hasattr(macho_instance, 'header') or not macho_instance.header:
    #             macho_instance.parse()
            
    #         # Get human-readable values using machofile's formatting methods
    #         magic_str = macho_instance.format_magic_value(header.get('magic', 0))
            
    #         # Simple CPU type mapping since CPU_TYPE_MAP is not exposed
    #         cputype = header.get('cputype', 0)
    #         if cputype == 0x7:
    #             cputype_str = "x86"
    #         elif cputype == 0x1000007:
    #             cputype_str = "x86_64"
    #         elif cputype == 0xC:
    #             cputype_str = "ARM"
    #         elif cputype == 0x100000C:
    #             cputype_str = "ARM 64-bit"
    #         else:
    #             cputype_str = str(cputype)
            
    #         cpusubtype_str = macho_instance.decode_cpusubtype(header.get('cputype', 0), header.get('cpusubtype', 0))
    #         filetype_str = macho_instance.format_file_type(header.get('filetype', 0))
    #         flags_str = macho_instance.decode_flags(header.get('flags', 0))
            
    #         return {
    #             'raw': {
    #                 'magic': header.get('magic', 0),
    #                 'cputype': header.get('cputype', 0),
    #                 'cpusubtype': header.get('cpusubtype', 0),
    #                 'filetype': header.get('filetype', 0),
    #                 'flags': header.get('flags', 0),
    #             },
    #             'formatted': {
    #                 'magic_str': magic_str,
    #                 'cputype_str': cputype_str,
    #                 'cpusubtype_str': cpusubtype_str,
    #                 'filetype_str': filetype_str,
    #                 'flags_str': flags_str,
    #             }
    #         }
    #     except Exception as e:
    #         self.log.error(f"Error formatting header values: {e}")
    #         return None