Rajeev Gandhi

51 papers A 7B 6C 3Misc 6Journal 5Unranked 24
YearRankTypeTitle / Venue / Authors
2017 A conf
ICDCS
Utsav Drolia, Katherine Guo, Jiaqi Tan, Rajeev Gandhi, Priya Narasimhan
2017 conf
PerCom Workshops
Utsav Drolia, Katherine Guo, Jiaqi Tan, Rajeev Gandhi, Priya Narasimhan
2016 B conf
APLAS
Jiaqi Tan, Hui Jun Tay, Rajeev Gandhi, Priya Narasimhan
2016 Misc conf
EMSOFT
Jiaqi Tan, Hui Jun Tay, Utsav Drolia, Rajeev Gandhi, Priya Narasimhan
2016 Misc conf
SEC
Jiaqi Tan, Rajeev Gandhi, Priya Narasimhan
2016 Misc conf
SEC
Utsav Drolia, Katherine Guo, Rajeev Gandhi, Priya Narasimhan
2016 conf
MobiSys (Companion Volume)
Utsav Drolia, Katherine Guo, Rajeev Gandhi, Priya Narasimhan
2016 C conf
WoWMoM
Nathan D. Mickulicz, Utsav Drolia, Priya Narasimhan, Rajeev Gandhi
2015 conf
VSTTE
Jiaqi Tan, Hui Jun Tay, Rajeev Gandhi, Priya Narasimhan
2015 conf
MobiArch
Utsav Drolia, Nathan D. Mickulicz, Rajeev Gandhi, Priya Narasimhan
2015 conf
S3@MobiCom
Utsav Drolia, Nathan D. Mickulicz, Rajeev Gandhi, Priya Narasimhan
2015 B conf
GPCE
Priya Narasimhan, Utsav Drolia, Jiaqi Tan, Nathan D. Mickulicz, Rajeev Gandhi
2015 conf
BigDataService
Nathan D. Mickulicz, Rolando Martins, Priya Narasimhan, Rajeev Gandhi
2014 conf
HCI (10)
Vinod Sharma, Kunal Mankodiya, Fernando De la Torre, Ada Zhang, Neal Ryan, Thanh G. N. Ton, Rajeev Gandhi, Samay Jain
2014 C conf
CloudCom
Jiaqi Tan, Rajeev Gandhi, Priya Narasimhan
2014 C conf
CloudCom
Jiaqi Tan, Utsav Drolia, Rolando Martins, Rajeev Gandhi, Priya Narasimhan
2014 conf
WISEC
Jiaqi Tan, Utsav Drolia, Rolando Martins, Rajeev Gandhi, Priya Narasimhan
2013 conf
CAMAD
Bruno Sousa, Ricardo Santos, Marília Curado, Soila M. Pertet, Rajeev Gandhi, Carlos Silva, Kostas Pentikousis
2013 A conf
Middleware
Rolando Martins, Rajeev Gandhi, Priya Narasimhan, Soila M. Pertet, António Casimiro, Diego Kreutz, Paulo Veríssimo
2013 conf
HCI (11)
Kunal Mankodiya, Rolando Martins, Jonathan Francis, Elmer Garduño, Rajeev Gandhi, Priya Narasimhan
2013 conf
ROSE
Jonathan Francis, Utsav Drolia, Kunal Mankodiya, Rolando Martins, Rajeev Gandhi, Priya Narasimhan
2013 J jnl
ACM SIGOPS Oper. Syst. Rev.
Chengwei Wang, Soila Kavulya, Jiaqi Tan, Liting Hu, Mahendra Kutare, Michael P. Kasick, Karsten Schwan, Priya Narasimhan, Rajeev Gandhi
2013 conf
UIC/ATC
Utsav Drolia, Rolando Martins, Jiaqi Tan, Ankit Chheda, Monil Mukesh Sanghavi, Rajeev Gandhi, Priya Narasimhan
2013 J jnl
login Usenix Mag.
Elmer Garduño, Soila Kavulya, Jiaqi Tan, Rajeev Gandhi, Priya Narasimhan
2013 conf
ICAC
Nathan D. Mickulicz, Priya Narasimhan, Rajeev Gandhi
2013 conf
UIC/ATC
Kunal Mankodiya, Vinod Sharma, Rolando Martins, Ishan Pande, Samay Jain, Neal Ryan, Rajeev Gandhi
2013 conf
LISA
Nathan D. Mickulicz, Priya Narasimhan, Rajeev Gandhi
2012 conf
S-CUBE
Kunal Mankodiya, Rajeev Gandhi, Priya Narasimhan
2012 A conf
DSN
Soila Kavulya, Scott Daniels, Kaustubh R. Joshi, Matti A. Hiltunen, Rajeev Gandhi, Priya Narasimhan
2012 conf
DSN Workshops
António Casimiro, Paulo Veríssimo, Diego Kreutz, Filipe Araújo, Raul Barbosa, Samuel Neves, Bruno Sousa, Marília Curado, Carlos Silva, Rajeev Gandhi, Priya Narasimhan
2012 conf
LISA
Elmer Garduño, Soila Kavulya, Jiaqi Tan, Rajeev Gandhi, Priya Narasimhan
2011 J jnl
ACM SIGOPS Oper. Syst. Rev.
Soila Kavulya, Kaustubh R. Joshi, Matti A. Hiltunen, Scott Daniels, Rajeev Gandhi, Priya Narasimhan
2010 B conf
CCGRID
Soila Kavulya, Jiaqi Tan, Rajeev Gandhi, Priya Narasimhan
2010 conf
HotDep
Michael P. Kasick, Rajeev Gandhi, Priya Narasimhan
2010 A conf
FAST
Michael P. Kasick, Jiaqi Tan, Rajeev Gandhi, Priya Narasimhan
2010 B conf
NOMS
Jiaqi Tan, Xinghao Pan, Eugene Marinelli, Soila Kavulya, Rajeev Gandhi, Priya Narasimhan
2010 A conf
ICDCS
Jiaqi Tan, Soila Kavulya, Rajeev Gandhi, Priya Narasimhan
2009 B conf
WADS
Keith A. Bare, Soila Kavulya, Jiaqi Tan, Xinghao Pan, Eugene Marinelli, Michael P. Kasick, Rajeev Gandhi, Priya Narasimhan
2009 J jnl
SIGMETRICS Perform. Evaluation Rev.
Xinghao Pan, Jiaqi Tan, Soila Kavulya, Rajeev Gandhi, Priya Narasimhan
2009 conf
HotCloud
Jiaqi Tan, Xinghao Pan, Soila Kavulya, Rajeev Gandhi, Priya Narasimhan
2009 Misc conf
CASES
Priya Narasimhan, Rajeev Gandhi, Dan Rossi
2008 B conf
SRDS
Soila Kavulya, Rajeev Gandhi, Priya Narasimhan
2008 conf
WASL
Jiaqi Tan, Xinghao Pan, Soila Kavulya, Rajeev Gandhi, Priya Narasimhan
2007 A conf
RTSS
Donnie H. Kim, Rajeev Gandhi, Priya Narasimhan
2007 conf
ICDCS Workshops
Donnie H. Kim, Rajeev Gandhi, Priya Narasimhan
2006 A conf
ICDCS
Patrick E. Lanigan, Rajeev Gandhi, Priya Narasimhan
2005 Misc conf
SenSys
Patrick E. Lanigan, Rajeev Gandhi, Priya Narasimhan
2005 J jnl
ACM Trans. Embed. Comput. Syst.
Philip Koopman, Howie Choset, Rajeev Gandhi, Bruce H. Krogh, Diana Marculescu, Priya Narasimhan, JoAnn M. Paul, Ragunathan Rajkumar, Daniel P. Siewiorek, Asim Smailagic, Peter Steenkiste, Donald E. Thomas, Chenxi Wang
2003 conf
WORDS Fall
Rajeev Gandhi
2001 Misc conf
ICASSP
Rajeev Gandhi, Sanjit K. Mitra
1998 conf
ICECS
Rajeev Gandhi, Sanjit K. Mitra
redb/extractors/extractor.py
← Index redb/extractors/extractor.py python
import hashlib
import inspect
from abc import ABCMeta, abstractmethod
from dataclasses import asdict
from functools import cached_property
from datetime import datetime, timezone
import math
from typing import Counter, List, Optional, Dict, Any, Tuple
from redb import settings
from redb.models.dataclasses import Hash
from .database_exporters import DatabaseExporter, ElasticsearchExporter, ClickHouseExporter, PrintExporter

class Extractor(metaclass=ABCMeta):
    def __init__(
        self,
        filepath: str,
        log: Any,
        exporters: Optional[List[DatabaseExporter]] = None,
        index_prefix: Optional[str] = None,
        source: Optional[str] = None,
        elastic_index: Optional[str] = None,
        known_benign: bool = False,
        known_malicious: bool = False,
        precomputed_hashes: Optional[Dict[str, str]] = None,
    ):
        self.log = log
        self.log.debug(f"Creating {self.__class__.__name__}")
        self.filepath = filepath
        self.source = source
        self.exporters = exporters or []
        self.index_prefix = index_prefix if index_prefix else settings.ELASTIC_BINARIES_COLLECTION
        self.elastic_index = self.index_prefix + (elastic_index if elastic_index else "")
        self.known_benign = known_benign
        self.known_malicious = known_malicious

        # Use precomputed hashes if provided (e.g., from machofile, pefile)
        # Otherwise compute them from binary
        if precomputed_hashes:
            self.md5 = precomputed_hashes.get('md5') or precomputed_hashes.get('MD5')
            self.sha1 = precomputed_hashes.get('sha1') or precomputed_hashes.get('SHA1')
            self.sha256 = precomputed_hashes.get('sha256') or precomputed_hashes.get('SHA256')
        else:
            self.md5 = hashlib.md5(self.binary).hexdigest()
            self.sha1 = hashlib.sha1(self.binary).hexdigest()
            self.sha256 = hashlib.sha256(self.binary).hexdigest()
        self.hash = Hash(self.md5, self.sha1, self.sha256)

    @cached_property
    def binary(self):
        with open(self.filepath, "rb") as f:
            data = f.read()
        return data

    @property
    @abstractmethod
    def tag(self):
        pass

    @abstractmethod
    def extract(self):
        """
        this method defines the extracted data
        """

    @staticmethod
    def process_binary_string(s):
        # Remove \x00 padding
        s = s.rstrip(b"\x00")

        # Check if there are any non-printable characters
        has_non_printable = any(byte < 32 or byte > 126 for byte in s)

        if not has_non_printable:
            # If all characters are printable, decode the string
            return s.decode()
        else:
            # If there are non-printable characters, represent them as \xDD
            return "".join(
                [
                    f"\\x{byte:02x}" if byte < 32 or byte > 126 else chr(byte)
                    for byte in s
                ]
            )

    @staticmethod
    def remove_non_utf8(binary_string):
        decoded = b""
        for i in range(len(binary_string)):
            try:
                # Try to decode each byte
                char = binary_string[i : i + 1].decode("utf-8")
                decoded += char.encode("utf-8")
            except UnicodeDecodeError:
                # Skip this byte if it can't be decoded
                continue
        return decoded

    def calculate_entropy(self, data):
        """Calculate the entropy of a chunk of data.
        Based on pefile.SectionStructure.entropy_H.
        """
        # self.log.debug(inspect.currentframe().f_code.co_name)
        if not data:
            return 0.0

        if type(data) == str:
            counts = Counter(data)
            frequencies = ((i / len(data)) for i in counts.values())
            return - sum(f * math.log(f, 2) for f in frequencies)
        else:
            occurences = Counter(bytearray(data))
            entropy = 0
            for x in occurences.values():
                p_x = float(x) / len(data)
                entropy -= p_x * math.log(p_x, 2)
            return entropy

    @abstractmethod
    def prepare_export_data(self, exporter_type: str) -> Tuple[List[Any], List[str], List[str]]:
        """
        Prepare data for specific export type
        Returns:
            Tuple containing:
            - data: List of values to insert
            - column_names: List of column names
            - column_type_names: List of column types
        """
        pass

    def export_data(self):
        """Export data to all configured exporters

        Returns:
            True: Export succeeded
            False: Export failed (actual error)
            None: No data to export (not an error, e.g., no overlay, no signature)
        """
        self.log.debug(inspect.currentframe().f_code.co_name)
        success = True
        extracted_data = self.extract()

        if extracted_data is None:
            self.log.debug("extract() returned None, skipping export")
            return None  # No data to export, not a failure
            
        for exporter in self.exporters:
            if isinstance(exporter, PrintExporter):
                # For PrintExporter, we pass the extracted data directly
                success &= exporter.export(extracted_data)
            else:
                # Get the data prepared for this specific exporter type
                export_data = self.prepare_export_data(exporter.__class__.__name__)
                
                if export_data is None:
                    self.log.debug(f"prepare_export_data returned None for {exporter.__class__.__name__}")
                    return False
                
                if isinstance(exporter, ElasticsearchExporter):
                    success &= exporter.export(
                        export_data,
                        index=self.elastic_index,
                        tag=self.tag(),
                        hashes=asdict(self.hash),
                        known_benign=self.known_benign,
                        known_malicious=self.known_malicious
                    )
                elif isinstance(exporter, ClickHouseExporter):
                    # For ClickHouse, we need to pass the table name and the prepared data
                    success &= exporter.export(
                        export_data,
                        table=self.get_clickhouse_table(),
                        # Add these parameters explicitly
                        column_names=export_data[1] if isinstance(export_data, tuple) else None,
                        column_type_names=export_data[2] if isinstance(export_data, tuple) else None
                    )
                
        return success

    @abstractmethod
    def get_clickhouse_table(self) -> str:
        """Return the appropriate ClickHouse table name"""
        pass
    # def export_to_elastic(self, list_of_dataclasses, tag=None):
    #     self.log.debug(inspect.currentframe().f_code.co_name)

    #     if not isinstance(list_of_dataclasses, list):
    #         self.log.error("Called export_to_elastic wrongly")

    #     now_t = datetime.now(timezone.utc).strftime("%Y-%m-%d %H:%M:%S")
    #     for dataclass_ in list_of_dataclasses:
    #         if not settings.ELASTIC_CLIENT.ping():
    #             self.log.error("[CONNECTION ERROR] ping to elastic failed")
            
    #         # Convert dataclass to dict and filter out None values
    #         document = {k: v for k, v in asdict(dataclass_).items() if v is not None}
            
    #         if tag:
    #             document["tag"] = [tag, self.tag()]
    #         else:
    #             document["tag"] = self.tag()
    #         hashes = asdict(self.hash)
    #         document |= hashes
    #         document["timestamp_utc"] = now_t
    #         # document["source"] = self.source
    #         document["known_benign"] = self.known_benign
    #         document["known_malicious"] = self.known_malicious

    #         if "_id" in document:
    #             tmp_id = document.pop("_id") + document["sha256"]
    #             _id = hashlib.sha256(tmp_id.encode()).hexdigest()
    #         else:
    #             _id = document["sha256"]

    #         # self.log.debug(f"[DEBUG] about to export {type(document)} {document}")
    #         try:
    #             doc_dump = json.dumps(document)
    #         except TypeError as e:
    #             self.log.error(
    #                 f"Failed export of document. " f"full document: {document}"
    #             )
    #             raise e

    #         # body={"doc": doc_dump,
    #         #       "doc_as_upsert": True  # Create the document if it doesn't exist
    #         # }

    #         # Check if the index exists, and create it if it doesn't
    #         # if not settings.ELASTIC_CLIENT.indices.exists(index=self.elastic_index):
    #         #     settings.ELASTIC_CLIENT.indices.create(index=self.elastic_index)
    #         # pprint(doc_dump) #DEBUG
    #         # print("[DEBUG] _id: " + _id)
    #         # print("[DEBUG] index: " + self.elastic_index)
    #         settings.ELASTIC_CLIENT.index(
    #             index=self.elastic_index, id=_id, document=doc_dump
    #         )