Rakesh B. Bobba

69 papers A* 6A 9B 5C 2Journal 27Unranked 17
YearRankTypeTitle / Venue / Authors
2025 A conf
ACSAC
Akshith Gunasekaran, Gabriel Ritter, Rakesh B. Bobba
2025 C conf
SECRYPT
Mahsa Saeidi, Sai Sree Laya Chukkapalli, Anita Sarma, Rakesh B. Bobba
2025 conf
CPSIOTSEC@CCS
Gabriel Ritter, Rakesh B. Bobba
2022 J jnl
Proc. Priv. Enhancing Technol.
Arezoo Rajabi, Mahdieh Abbasi, Rakesh B. Bobba, Kimia Tajik
2022 J jnl
ACM Trans. Cyber Phys. Syst.
Monowar Hasan, Sibin Mohan, Rakesh B. Bobba, Rodolfo Pellizzoni
2022 J jnl
Proc. Priv. Enhancing Technol.
Mahsa Saeidi, McKenzie Calvert, Audrey Au, Anita Sarma, Rakesh B. Bobba
2021 J jnl
Proc. Priv. Enhancing Technol.
Arezoo Rajabi, Rakesh B. Bobba, Mike Rosulek, Charles V. Wright, Wu-chi Feng
2021 J jnl
IEEE Trans. Smart Grid
Arezoo Rajabi, Rakesh B. Bobba
2021 A* conf
INFOCOM
Ashish Kashinath, Monowar Hasan, Rakesh Kumar, Sibin Mohan, Rakesh B. Bobba, Smruti Padhy
2020 J jnl
CoRR
Arezoo Rajabi, Rakesh B. Bobba
2020 J jnl
CoRR
Mahsa Saeidi, McKenzie Calvert, Audrey Au, Anita Sarma, Rakesh B. Bobba
2020 conf
GLOBECOM (Workshops)
Ashish Kashinath, Monowar Hasan, Sibin Mohan, Rakesh B. Bobba, Radhika Mittal
2020 J jnl
CoRR
Chien-Ying Chen, Sibin Mohan, Rodolfo Pellizzoni, Rakesh B. Bobba
2020 A conf
DATE
Monowar Hasan, Sibin Mohan, Rodolfo Pellizzoni, Rakesh B. Bobba
2020 conf
Canadian AI
Mahdieh Abbasi, Arezoo Rajabi, Christian Gagné, Rakesh B. Bobba
2020 J jnl
CoRR
Mahdieh Abbasi, Arezoo Rajabi, Christian Gagné, Rakesh B. Bobba
2020 A conf
ECAI
Mahdieh Abbasi, Changjian Shui, Arezoo Rajabi, Christian Gagné, Rakesh B. Bobba
2019 A conf
RTAS
Chien-Ying Chen, Sibin Mohan, Rodolfo Pellizzoni, Rakesh B. Bobba, Negar Kiyavash
2019 A* conf
NDSS
Kimia Tajik, Akshith Gunasekaran, Rhea Dutta, Brandon Ellis, Rakesh B. Bobba, Mike Rosulek, Charles V. Wright, Wu-chi Feng
2019 J jnl
IACR Cryptol. ePrint Arch.
Kimia Tajik, Akshith Gunasekaran, Rhea Dutta, Brandon Ellis, Rakesh B. Bobba, Mike Rosulek, Charles V. Wright, Wu-chi Feng
2019 conf
SmartGridComm
Arezoo Rajabi, Rakesh B. Bobba
2019 J jnl
CoRR
Hsuan-Chi Kuo, Akshith Gunasekaran, Yeongjin Jang, Sibin Mohan, Rakesh B. Bobba, David Lie, Jesse Walker
2019 J jnl
CoRR
Monowar Hasan, Sibin Mohan, Rodolfo Pellizzoni, Rakesh B. Bobba
2018 A conf
DATE
Monowar Hasan, Sibin Mohan, Rodolfo Pellizzoni, Rakesh B. Bobba
2018 J jnl
CoRR
Mahdieh Abbasi, Arezoo Rajabi, Azadeh Sadat Mozafari, Rakesh B. Bobba, Christian Gagné
2018 conf
CPS-SPC@CCS
Vedanth Narayanan, Rakesh B. Bobba
2018 J jnl
CoRR
Chien-Ying Chen, Sibin Mohan, Rakesh B. Bobba, Rodolfo Pellizzoni, Negar Kiyavash
2018 B conf
IC2E
Read Sprabery, Konstantin Evchenko, Abhilash Raj, Rakesh B. Bobba, Sibin Mohan, Roy H. Campbell
2018 J jnl
CoRR
Arezoo Rajabi, Mahdieh Abbasi, Christian Gagné, Rakesh B. Bobba
2017 J jnl
CoRR
Monowar Hasan, Sibin Mohan, Rodolfo Pellizzoni, Rakesh B. Bobba
2017 J jnl
CoRR
Read Sprabery, Konstantin Evchenko, Abhilash Raj, Rakesh B. Bobba, Sibin Mohan, Roy H. Campbell
2017 J jnl
CoRR
Chien-Ying Chen, AmirEmad Ghassami, Sibin Mohan, Negar Kiyavash, Rakesh B. Bobba, Rodolfo Pellizzoni, Man-Ki Yoon
2017 conf
MPS@CCS
Byron Marohn, Charles V. Wright, Wu-chi Feng, Mike Rosulek, Rakesh B. Bobba
2017 J jnl
IACR Cryptol. ePrint Arch.
Byron Marohn, Charles V. Wright, Wu-chi Feng, Mike Rosulek, Rakesh B. Bobba
2017 A* conf
CCS
Rakesh B. Bobba, Awais Rashid
2017 B conf
ECRTS
Monowar Hasan, Sibin Mohan, Rodolfo Pellizzoni, Rakesh B. Bobba
2017 J jnl
CoRR
Monowar Hasan, Sibin Mohan, Rodolfo Pellizzoni, Rakesh B. Bobba
2017 J jnl
CoRR
Rakesh Kumar, Monowar Hasan, Smruti Padhy, Konstantin Evchenko, Lavanya Piramanayagam, Sibin Mohan, Rakesh B. Bobba
2017 A conf
RTSS
Rakesh Kumar, Monowar Hasan, Smruti Padhy, Konstantin Evchenko, Lavanya Piramanayagam, Sibin Mohan, Rakesh B. Bobba
2017 A* conf
CHI
Jun Ho Huh, Hyoungshick Kim, Swathi S. V. P. Rayala, Rakesh B. Bobba, Konstantin Beznosov
2017 ed.
CPS-SPC@CCS
Bhavani Thuraisingham, Rakesh B. Bobba, Awais Rashid
2017 B conf
IC2E
Read Sprabery, Zachary John Estrada, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Rakesh B. Bobba, Roy H. Campbell
2017 B conf
VL/HCC
Kim J. Kaaz, Alex Hoffer, Mahsa Saeidi, Anita Sarma, Rakesh B. Bobba
2016 conf
ICSS
Arezoo Rajabi, Rakesh B. Bobba
2016 conf
SmartGridComm
Gabriel A. Weaver, Kate Davis, Charles M. Davis, Edmond J. Rogers, Rakesh B. Bobba, Saman A. Zonouz, Robin Berthier, Peter W. Sauer, David M. Nicol
2016 A conf
RTSS
Monowar Hasan, Sibin Mohan, Rakesh B. Bobba, Rodolfo Pellizzoni
2016 J jnl
CoRR
Monowar Hasan, Sibin Mohan, Rakesh B. Bobba, Rodolfo Pellizzoni
2016 J jnl
Real Time Syst.
Sibin Mohan, Man-Ki Yoon, Rodolfo Pellizzoni, Rakesh B. Bobba
2016 J jnl
IEEE Internet Comput.
Jun Ho Huh, Rakesh B. Bobba, Tom Markham, David M. Nicol, Julie Hull, Alexander Chernoguzov, Himanshu Khurana, Kevin Staggs, Jingwei Huang
2016 ed.
CPS-SPC@CCS
Edgar R. Weippl, Stefan Katzenbeisser, Mathias Payer, Stefan Mangard, Alvaro A. Cárdenas, Rakesh B. Bobba
2016 A* conf
CCS
Alvaro A. Cárdenas, Rakesh B. Bobba
2015 J jnl
IEEE Trans. Smart Grid
Katherine R. Davis, Charles M. Davis, Saman A. Zonouz, Rakesh B. Bobba, Robin Berthier, Luis Garcia, Peter W. Sauer
2015 A* conf
CCS
Roshan K. Thomas, Alvaro A. Cárdenas, Rakesh B. Bobba
2015 conf
CNS
Weijie Liu, Rakesh B. Bobba, Sibin Mohan, Roy H. Campbell
2015 A conf
SOUPS
Jun Ho Huh, Hyoungshick Kim, Rakesh B. Bobba, Masooda N. Bashir, Konstantin Beznosov
2014 J jnl
IEEE Trans. Smart Grid
Alvaro A. Cárdenas, Robin Berthier, Rakesh B. Bobba, Jun Ho Huh, Jorjeta G. Jetcheva, David Grochocki, William H. Sanders
2014 conf
SmartGridComm
Tawfeeq A. Shawly, Jun Liu, Nathan Burow, Saurabh Bagchi, Robin Berthier, Rakesh B. Bobba
2014 conf
MTD@CCS
Mohammad Ashiqur Rahman, Ehab Al-Shaer, Rakesh B. Bobba
2014 conf
SmartGridComm
Robin Berthier, David I. Urbina, Alvaro A. Cárdenas, Michael Guerrero, Ulrich Herberg, Jorjeta G. Jetcheva, Daisuke Mashima, Jun Ho Huh, Rakesh B. Bobba
2014 J jnl
IEEE Trans. Smart Grid
Saman A. Zonouz, Charles M. Davis, Katherine R. Davis, Robin Berthier, Rakesh B. Bobba, William H. Sanders
2014 C conf
ACC
André Teixeira, György Dán, Henrik Sandberg, Robin Berthier, Rakesh B. Bobba, Alfonso Valdes
2013 conf
HotCloud
György Dán, Rakesh B. Bobba, George Gross, Roy H. Campbell
2013 conf
SmartGridComm
Ognjen Vukovic, György Dán, Rakesh B. Bobba
2013 B conf
CCGRID
Stephen Skeirik, Rakesh B. Bobba, José Meseguer
2013 conf
SmartGridComm
Muhammad Salman Malik, Robin Berthier, Rakesh B. Bobba, Roy H. Campbell, William H. Sanders
2013 conf
SmartGridComm
William Niemira, Rakesh B. Bobba, Peter W. Sauer, William H. Sanders
2013 conf
SmartGridComm
Robin Berthier, Jorjeta G. Jetcheva, Daisuke Mashima, Jun Ho Huh, David Grochocki, Rakesh B. Bobba, Alvaro A. Cárdenas, William H. Sanders
2013 A conf
DSN
Muhammad Salman Malik, Mirko Montanari, Jun Ho Huh, Rakesh B. Bobba, Roy H. Campbell
2009
Rakesh B. Bobba
redb/extractors/decompiler/bninja/analysis/disassembly.py
← Index redb/extractors/decompiler/bninja/analysis/disassembly.py python
import re
import time

import binaryninja
from binaryninja.enums import (
    InstructionTextTokenType,
)

# Support both package and standalone imports
try:
    from ..function_type import FunctionTypeAnalysis
    from ..utils.hashes import calculate_sha256
except ImportError:
    # Fallback to absolute imports (for multiprocessing spawned processes)
    from redb.extractors.decompiler.bninja.function_type import FunctionTypeAnalysis
    from redb.extractors.decompiler.bninja.utils.hashes import calculate_sha256


class DisassemblyAnalysis:
    INVALID_STACK_SIZE = -1

    def __init__(self, arch, function, bv, logger):
        self.arch = arch
        self.function = function
        self.bv = bv
        self.logger = logger
        if self.function is not None and hasattr(self.function, "instructions"):
            self.instructions = self.function.instructions
        else:
            self.instructions = []
        self.errors = []
        return

    def log_error(
        self, message, function_name, address, exception=None, error_location="unknown"
    ):
        """Log an error during processing."""
        error_msg = f"Error in function {function_name} at {address}: {message}"
        if exception:
            error_msg += f" - {str(exception)}"
        self.logger.error(error_msg)

        # Add to errors list
        error = {
            "function_name": function_name,
            "function_address": str(address),
            "error_location": error_location,
            "error_message": message,
            "error_details": str(exception) if exception else "",
            "error_type": type(exception).__name__ if exception else "Unknown",
            "timestamp": int(time.time() * 1000),
        }
        self.errors.append(error)

    def get_json(self):
        try:
            # Build disassembly string and normalized versions
            disassembly_builder = [[], []]  # Address and instruction text

            # Create a dictionary mapping addresses to instruction tokens
            instr_tokens_by_addr = {}
            for instr_tokens, addr in self.instructions:
                instr_tokens_by_addr[addr] = instr_tokens

            addresses = sorted(instr_tokens_by_addr.keys())
            for address in addresses:
                # Original disassembly with addresses
                # instr_tokens, address = instruction
                instr_tokens = instr_tokens_by_addr[address]
                disassembly_builder[0].append(address)
                disassembly_builder[1].append("".join(map(str, instr_tokens)))

            # Join with newlines
            disassembly_str = "\n".join(disassembly_builder[1])
            disassembly_with_addresses = "\n".join(
                f"{hex(address)}: {instr_text}"
                for address, instr_text in zip(
                    disassembly_builder[0], disassembly_builder[1], strict=False
                )
            )

            disassembly_json = {
                "disassembled_function_hash": calculate_sha256(disassembly_str),
                "disassembled_function": disassembly_with_addresses,
                "disassembled_function_no_addresses": disassembly_str,
                "disassembled_function_name": self.function.name,
                "disassembled_function_address": self.function.start,
                "instructions_count": len(instr_tokens_by_addr.keys()),
                "function_type": FunctionTypeAnalysis(self.function)
                .get_function_type()
                .name,
            }

            # Add additional metrics
            type_frequencies = self.collect_instruction_types()
            disassembly_json["instructions_types"] = list(type_frequencies.keys())
            disassembly_json["control_flow_count"] = (
                self.count_control_flow_instructions()
            )
            disassembly_json["memory_access_pattern"] = self.collect_memory_patterns()
            disassembly_json["register_usage"] = self.collect_register_usage()
            disassembly_json["data_references_count"] = self.count_data_references()
            disassembly_json["max_block_size"] = self.compute_max_block_size()
            disassembly_json["num_calls"] = self.compute_num_calls()
            disassembly_json["stack_size"] = self.estimate_stack_size()

            return disassembly_json, self.errors

        except Exception as e:
            self.log_error(
                "Failed to collect instruction types",
                self.function.name,
                self.function.start,
                e,
                "collect_instruction_types",
            )
            raise ValueError(e) from e

    def collect_instruction_types(self):
        """Collect instruction type frequencies from a function."""
        type_frequencies = {}

        try:
            # Iterate through all instructions in the function
            for instruction in self.instructions:
                instr_tokens = instruction[0]  # Get the instruction tokens

                # Extract the mnemonic from the instruction tokens
                mnemonic = None
                for token in instr_tokens:
                    if token.type == InstructionTextTokenType.InstructionToken:
                        mnemonic = token.text
                        break

                if not mnemonic:
                    continue

                # Use normalize_opcode to get standardized opcode
                normalized = self.normalize_opcode(mnemonic)

                # Get category from opcode_categories or use the instruction type directly
                category = self.arch.opcode_categories.get(normalized)
                if category:
                    self._increment_frequency(type_frequencies, category)

        except Exception as e:
            self.log_error(
                "Failed to collect instruction types",
                self.function.name,
                self.function.start,
                e,
                "collect_instruction_types",
            )

        return type_frequencies

    def normalize_opcode(self, opcode):
        return opcode.upper()

    def collect_memory_patterns(self):
        """Collect memory access patterns from a function."""
        patterns = []
        try:
            for instruction in self.instructions:
                instr_tokens = instruction[0]

                # We need to capture memory operands between BeginMemoryOperandToken and EndMemoryOperandToken
                in_memory_operand = False
                memory_operand_text = ""

                for token in instr_tokens:
                    if token.type == InstructionTextTokenType.BeginMemoryOperandToken:
                        in_memory_operand = True
                        memory_operand_text = ""
                    elif token.type == InstructionTextTokenType.EndMemoryOperandToken:
                        in_memory_operand = False

                        # Process the captured memory operand text
                        if memory_operand_text:
                            # Categorize memory access pattern
                            if (
                                "+" in memory_operand_text
                                and "*" in memory_operand_text
                            ):
                                if "MEM_SCALED_INDEX" not in patterns:
                                    patterns.append("MEM_SCALED_INDEX")
                            elif (
                                "+" in memory_operand_text or "-" in memory_operand_text
                            ):
                                if "MEM_BASE_OFFSET" not in patterns:
                                    patterns.append("MEM_BASE_OFFSET")
                            else:
                                if "MEM_DIRECT" not in patterns:
                                    patterns.append("MEM_DIRECT")

                            # Check for stack accesses
                            if any(
                                reg in memory_operand_text
                                for reg in ["SP", "BP", "ESP", "EBP", "RSP", "RBP"]
                            ):
                                if "MEM_STACK" not in patterns:
                                    patterns.append("MEM_STACK")
                            # Check for string operations
                            elif (
                                any(
                                    reg in memory_operand_text
                                    for reg in ["SI", "DI", "ESI", "EDI", "RSI", "RDI"]
                                )
                                and "MEM_STRING" not in patterns
                            ):
                                patterns.append("MEM_STRING")
                    elif in_memory_operand:
                        # Accumulate token text while inside a memory operand
                        memory_operand_text += token.text
        except Exception as e:
            self.log_error(
                "Failed to collect memory patterns",
                self.function.name,
                self.function.start,
                e,
                "collect_memory_patterns",
            )
        return patterns

    def collect_register_usage(self):
        """Collect register usage from a function."""
        registers = []
        try:
            # Define register groups we're interested in tracking
            register_groups = {
                "GPR": [
                    "RAX",
                    "RBX",
                    "RCX",
                    "RDX",
                    "R9",
                    "R10",
                    "R11",
                    "R12",
                    "R13",
                    "R14",
                    "R15",
                    "EAX",
                    "EBX",
                    "ECX",
                    "EDX",
                    "R9D",
                    "R10D",
                    "R11D",
                    "R12D",
                    "R13D",
                    "R14D",
                    "AX",
                    "BX",
                    "CX",
                    "DX",
                ],
                "GPR_INDEX": ["RSI", "RDI", "ESI", "EDI", "SI", "DI"],
                "GPR_STACK": ["RSP", "RBP", "ESP", "EBP", "SP", "BP"],
                "SIMD": ["XMM", "YMM", "ZMM"],
                "FPU": ["ST", "ST0", "ST1", "ST2", "ST3", "ST4", "ST5", "ST6", "ST7"],
                "FLAGS": ["FLAGS", "EFLAGS", "RFLAGS"],
                "CONTROL_REGISTER": ["CR0", "CR2", "CR3", "CR4", "CR8"],
                "DEBUG_REGISTER": ["DR0", "DR1", "DR2", "DR3", "DR6", "DR7"],
            }

            # Extract registers from instructions
            for instruction in self.instructions:
                instr_tokens = instruction[0]
                for token in instr_tokens:
                    if token.type == InstructionTextTokenType.RegisterToken:
                        reg = token.text.upper()
                        # Check which group this register belongs to
                        for group, regs in register_groups.items():
                            # if any(r in reg for r in regs) or any(reg.startswith(r) for r in regs):
                            if any(reg == r or reg.startswith(r) for r in regs):
                                if group not in registers:
                                    registers.append(group)
                                break
        except Exception as e:
            self.log_error(
                "Failed to collect register usage",
                self.function.name,
                self.function.start,
                e,
                "collect_register_usage",
            )
        return registers

    def count_data_references(self):
        """Count the number of data references in a function."""
        count = 0
        try:

            if self.function.mlil is None:
                return 0

            for block in self.function.mlil:
                for instr in block:
                    instr_str = str(instr)
                    logged = False
                    src = None

                    # Check for constant dereferencing or symbolic refs
                    if hasattr(instr, "src"):
                        src = instr.src
                        if isinstance(
                            src,
                            (
                                binaryninja.mediumlevelil.MediumLevelILConstPtr,
                                binaryninja.mediumlevelil.MediumLevelILConst,
                            ),
                        ):
                            count += 1
                            logged = True

                    # Check full string for hardcoded addresses or symbol-like tokens
                    if re.search(r"\b0x[0-9A-Fa-f]{3,}\b", instr_str) and not logged:
                        count += 1
                        logged = True

                    if "_" in instr_str and not logged:
                        count += 1
                        logged = True

                    # Only check for MediumLevelILConstPtr if src exists
                    if src is not None and isinstance(
                        src, binaryninja.mediumlevelil.MediumLevelILConstPtr
                    ):
                        addr = src.constant
                        # Check if address is in data sections
                        segment = self.bv.get_segment_at(addr)
                        if segment and segment.writable:
                            # print(f"[{function.name}] Matched data section reference in: {instr_str}")
                            count += 1
                            logged = True
        except Exception as e:
            self.logger.warning(
                f"Failed to use MLIL for counting data references in {self.function.name} at {self.function.start}: {e}"
            )
        return count

    def compute_max_block_size(self):
        """Compute the maximum basic block size in a function."""
        max_size = 0
        if self.function is None:
            return 0

        for block in self.function.basic_blocks:
            try:
                # Count instructions in this block using the direct length approach
                # This avoids UTF-8 decoding issues entirely
                block_size = block.instruction_count
                max_size = max(max_size, block_size)
            except Exception as e:
                self.log_error(
                    f"[HandledError] computing max block size: {e}",
                    self.function.name,
                    self.function.start,
                    e,
                    "compute_max_block_size",
                )
        return max_size

    def count_control_flow_instructions(self):
        """Count the number of control flow instructions in a function."""
        count = 0
        try:
            for instruction in self.instructions:
                instr_tokens = instruction[0]
                if self.arch.is_control_flow_instruction(instr_tokens):
                    count += 1
        except Exception as e:
            self.log_error(
                "Failed to count control flow instructions",
                self.function.name,
                self.function.start,
                e,
                "count_control_flow_instructions",
            )
        return count

    def compute_num_calls(self) -> int:
        """Compute the number of call instructions in a function."""
        num_calls = 0
        try:
            for instruction in self.instructions:
                instr_tokens = instruction[0]
                # Extract the mnemonic
                for token in instr_tokens:
                    if token.type == InstructionTextTokenType.InstructionToken:
                        if token.text.upper() == "CALL":
                            num_calls += 1
                        break
        except Exception as e:
            self.log_error(
                "Failed to compute number of calls",
                self.function.name,
                self.function.start,
                e,
                "compute_num_calls",
            )
        return num_calls

    def _increment_frequency(self, frequencies, type_name):
        """Increment the frequency count for an instruction type."""
        if type_name in frequencies:
            frequencies[type_name] += 1
        else:
            frequencies[type_name] = 1

    def estimate_stack_size(self):
        """Estimate the stack size used by a function."""
        try:
            # Binary Ninja provides a stack adjustment value for functions
            # Need to convert OffsetWithConfidence to a plain integer
            stack_adjust = self.function.stack_adjustment
            if hasattr(stack_adjust, "value"):  # Handle OffsetWithConfidence objects
                return stack_adjust.value
            return stack_adjust
        except Exception as e:
            self.log_error(
                "Failed to estimate stack size",
                self.function.name,
                self.function.start,
                e,
                "estimate_stack_size",
            )
            return self.INVALID_STACK_SIZE