Carlos Vega

38 papers C 4Misc 1Journal 26Unranked 7
YearRankTypeTitle / Venue / Authors
2026 J jnl
Data
Carlos Vega, Norberto Medina, Raquel León, Himar Fabelo, Alicia Martín, Gustavo M. Callicó
2026 conf
VISAPP (2)
Martin Rydlo, Max Verbers, Carlos Vega, Raquel León, Himar Fabelo, Gustav Burström, Alfonso Lagares, Jesús Morera, Gustavo M. Callicó, Francesca Manni, Svitlana Zinger
2025 conf
Agentic AI/CREATE/Clinical MLLMs@MICCAI
Beatriz Garcia Santa Cruz, Carlos Vega, Philip Santangelo, Venkata P. Satagopam
2025 C conf
DSD
Alvaro Falcon, Carlos Vega, Gustavo M. Callicó
2025 J jnl
npj Digit. Medicine
Rebecca Ting Jiin Loo, Lukas Pavelka, Graziella Mangone, Fouad Khoury, Marie Vidailhet, Jean-Christophe Corvol, Enrico Glaab, Geeta Acharya, Gloria Aguayo, Myriam Alexandre, Muhammad Ali, Wim Ammerlann, Giuseppe Arena, Michele Bassis, Roxane Batutu, Katy Beaumont, Sibylle Béchet, Guy Berchem, Alexandre Bisdorff, Ibrahim Boussaad, David Bouvier, Lorieza Castillo, Gessica Contesotto, Nancy De Bremaeker, Brian Dewitt, Nico Diederich, Rene Dondelinger, Nancy E. Ramia, Maria Fernanda Niño Uribe, Angelo Ferrari, Ana Festas Lopes, Katrin Frauenknecht, Joëlle Fritz, Carlos Gamio, Manon Gantenbein, Piotr Gawron, Laura Georges, Soumyabrata Ghosh, Marijus Giraitis, Martine Goergen, Elisa Gómez de Lope, Jérôme Graas, Mariella Graziano, Valentin Grouès, Anne Grünewald, Gaël Hammot, Anne-Marie Hanff, Linda Hansen, Michael Heneka, Estelle Henry, Margaux Henry, Sylvia Herbrink, Sascha Herzinger, Alexander Hundt, Nadine Jacoby, Sonja Jónsdóttir, Jochen Klucken, Olga Kofanova, Rejko Krüger, Pauline Lambert, Zied Landoulsi, Roseline Lentz, Victoria Lorentz, Tainá M. Marques, Guilherme Marques, Patricia Martins Conde, Patrick May, Deborah Mcintyre, Chouaib Mediouni, Francoise Meisch, Alexia Mendibide, Myriam Menster, Maura Minelli, Michel Mittelbronn, Saïda Mtimet, Maeva Munsch, Romain Nati, Ulf Nehrbass, Sarah Nickels, Beatrice Nicolai, Jean-Paul Nicolay, Fozia Noor, Clarissa P. C. Gomes, Sinthuja Pachchek, Claire Pauly, Laure Pauly, Magali Perquin, Achilleas Pexaras, Armin Rauschenberger, Rajesh Rawal, Dheeraj Reddy Bobbili, Lucie Remark, Ilsé Richard, Olivia Roland, Kirsten Roomp, Eduardo Rosales, Stefano Sapienza, Venkata P. Satagopam, Sabine Schmitz, Reinhard Schneider, Jens Schwamborn, Raquel Severino, Amir Sharify, Ruxandra Soare, Ekaterina Soboleva, Kate Sokolowska, Maud Theresine, Hermann Thien, Elodie Thiry, Johanna Trouet, Olena Tsurkalenko, Michel Vaillant, Carlos Vega, Liliana Vilas Boas, Paul Wilmes, Evi Wollscheid-Lengeling, Gelani Zelimkhanov, Marie-Alexandrine, Isabelle Arnulf, Samir Bekadar, Eve Benchetrit, Alexis Brice, Alizé Chalançon, Benoit Colsch, Florence Cormier-Dequaire, Virginie Czernecki, Bertrand Degos, Pauline Dodet, Carole Dongmo-Kenfack, Cécile Galléa, Rahul Gaurav, Manon Gomes, David Grabli, Marie-Odile Habert, Élodie Hainque, Farid Ichou, Jonas Ihle, Laetitia Jeancolas, Christelle Laganot, Louis-Laure Mariani, Mickaël Lé, Stéphane Lehéricy, Suzanne Lesage, Smaranda Leu Semenescu, Richard Levy, Valentine Maheo, Poornima Menon, Fanny Mochel, Vincent Perlbarg, Dijana Petrovska, Fanny Pineau, Nadya Pyatigorskaya, Sophie Rivaud-Péchoux, Sara Sambin, Julie Socha, Arthur Tenenhaus, Romain Valabrègue, Caroline Weill, Lydia Yahia-Cherif
2025 J jnl
Quantum
Carlos Vega, Alberto Muñoz de las Heras, Diego Porras, Alejandro González-Tudela
2024 C conf
DSD
Laura Quintana, Carlos Vega, Raquel León, Guillermo V. Socorro-Marrero, Samuel Ortega, Gustavo M. Callicó
2024 J jnl
Database J. Biol. Databases Curation
Carlos Vega, Marek Ostaszewski, Valentin Grouès, Reinhard Schneider, Venkata P. Satagopam
2024 J jnl
npj Digit. Medicine
Cyril Brzenczek, Quentin Klopfenstein, Tom Hähnel, Holger Fröhlich, Enrico Glaab, Geeta Acharya, Gloria Aguayo, Myriam Alexandre, Muhammad Ali, Wim Ammerlann, Giuseppe Arena, Michele Bassis, Roxane Batutu, Katy Beaumont, Sibylle Béchet, Guy Berchem, Alexandre Bisdorff, Ibrahim Boussaad, David Bouvier, Lorieza Castillo, Gessica Contesotto, Nancy De Bremaeker, Brian Dewitt, Nico Diederich, Rene Dondelinger, Nancy E. Ramia, Angelo Ferrari, Katrin Frauenknecht, Joëlle Fritz, Carlos Gamio, Manon Gantenbein, Piotr Gawron, Laura Georges, Soumyabrata Ghosh, Marijus Giraitis, Martine Goergen, Elisa Gómez de Lope, Jérôme Graas, Mariella Graziano, Valentin Grouès, Anne Grünewald, Gaël Hammot, Anne-Marie Hanff, Linda Hansen, Michael Heneka, Estelle Henry, Margaux Henry, Sylvia Herbrink, Sascha Herzinger, Alexander Hundt, Nadine Jacoby, Sonja Jónsdóttir, Jochen Klucken, Olga Kofanova, Rejko Krüger, Pauline Lambert, Zied Landoulsi, Roseline Lentz, Ana Festas Lopes, Victoria Lorentz, Tainá M. Marques, Guilherme Marques, Patricia Martins Conde, Patrick May, Deborah Mcintyre, Chouaib Mediouni, Francoise Meisch, Alexia Mendibide, Myriam Menster, Maura Minelli, Michel Mittelbronn, Saïda Mtimet, Maeva Munsch, Romain Nati, Ulf Nehrbass, Sarah Nickels, Beatrice Nicolai, Jean-Paul Nicolay, Maria Fernanda Niño Uribe, Fozia Noor, Clarissa P. C. Gomes, Sinthuja Pachchek, Claire Pauly, Laure Pauly, Lukas Pavelka, Magali Perquin, Achilleas Pexaras, Armin Rauschenberger, Rajesh Rawal, Dheeraj Reddy Bobbili, Lucie Remark, Ilsé Richard, Olivia Roland, Kirsten Roomp, Eduardo Rosales, Stefano Sapienza, Venkata P. Satagopam, Sabine Schmitz, Reinhard Schneider, Jens Schwamborn, Raquel Severino, Amir Sharify, Ruxandra Soare, Ekaterina Soboleva, Kate Sokolowska, Maud Theresine, Hermann Thien, Elodie Thiry, Rebecca Ting Jiin Loo, Johanna Trouet, Olena Tsurkalenko, Michel Vaillant, Carlos Vega, Liliana Vilas Boas, Paul Wilmes, Evi Wollscheid-Lengeling, Gelani Zelimkhanov
2024 C conf
DSD
Nerea Marquez-Suarez, Carlos Vega, Raquel León, Gustavo M. Callicó
2024 J jnl
J. Open Source Softw.
Andres R. Tejedor, Ignacio Sanchez-Burgos, Eduardo Sanz, Carlos Vega, Felipe J. Blas, Ruslan L. Davidchack, Nicodemo Di Pasquale, Jorge Ramirez, Jorge R. Espinosa
2023 J jnl
J. Medical Syst.
Carlos Vega, Reinhard Schneider, Venkata P. Satagopam
2023 conf
Digital and Computational Pathology
Carlos Vega, Laura Quintana, Samuel Ortega, Himar Fabelo, Esther Sauras, Noèlia Gallardo, Daniel Mata Cano, Marylène Lejeune, Carlos López, Gustavo M. Callicó
2022 J jnl
Data Sci.
Valentin Grouès, Carlos Vega, Venkata P. Satagopam
2022 C conf
DSD
Carlos Vega, Raquel León, Norberto Medina, Himar Fabelo, Samuel Ortega, Fran Balea, Aday García, Margarita Medina, Silvia De León, Alicia Martín, Gustavo M. Callicó
2022 conf
IWBBIO (2)
Carlos Vega, Miroslav Kratochvíl, Venkata P. Satagopam, Reinhard Schneider
2021 J jnl
IEEE Access
Carlos Vega
2021 conf
CIBB
Beatriz Garcia Santa Cruz, Carlos Vega, Frank Hertel
2019 J jnl
CoRR
Carlos Vega, Javier Aracil, Eduardo Magaña
2018 J jnl
CoRR
Carlos Vega, Eduardo Miravalls-Sierra, Guillermo Julián-Moreno, Jorge E. López de Vergara, Eduardo Magaña, Javier Aracil
2018 Misc conf
SoftCOM
Carlos Vega, Javier Aracil, Eduardo Magaña
2018 J jnl
Int. J. Netw. Manag.
Carlos Vega, Eduardo Miravalls-Sierra, Guillermo Julián-Moreno, Jorge E. López de Vergara, Eduardo Magaña, Javier Aracil
2017 conf
HPCC/SmartCity/DSS
Carlos Vega, Jose Fernando Zazo, Hugo Meyer, Ferad Zyulkyarov, Sergio López-Buedo, Javier Aracil
2017 J jnl
CoRR
Carlos Vega, Jose Fernando Zazo, Hugo Meyer, Ferad Zyulkyarov, Sergio López-Buedo, Javier Aracil
2017 J jnl
Netw. Protoc. Algorithms
Daniel Perdices, Jorge E. López de Vergara, Paula Roquero, Carlos Vega, Javier Aracil
2017 J jnl
CoRR
Carlos Vega, Paula Roquero, Rafael Leira, Iván González, Javier Aracil
2017 J jnl
J. Supercomput.
Carlos Vega, Paula Roquero, Rafael Leira, Iván González, Javier Aracil
2017 J jnl
CoRR
Carlos Vega, Paula Roquero, Javier Aracil
2017 J jnl
Comput. Networks
Carlos Vega, Paula Roquero, Javier Aracil
2013 J jnl
Comput. Phys. Commun.
Carl McBride, Eva G. Noya, Carlos Vega
2013 J jnl
Educ. Inf. Technol.
Carlos Vega, Camilo Jiménez, Jorge Villalobos
2012 conf
CSEDU (2)
Carlos Vega, Camilo Jiménez, Jorge Villalobos
2005 J jnl
Math. Comput. Model.
Sergio Sánchez, Regino Criado, Carlos Vega
2005 J jnl
Comput. Phys. Commun.
Carl McBride, Carlos Vega, Eduardo Sanz
1998 J jnl
SIAM Rev.
Andrés Bujosa, Regino Criado, Carlos Vega
1994 J jnl
Comput. Chem.
Carlos Vega, Santiago Lago
1989 J jnl
Kybernetika
Jirí Andel, María Gómez, Carlos Vega
1988 J jnl
Comput. Chem.
S. Lago, Carlos Vega
redb/extractors/decompiler/bninja/analysis/disassembly.py
← Index redb/extractors/decompiler/bninja/analysis/disassembly.py python
import re
import time

import binaryninja
from binaryninja.enums import (
    InstructionTextTokenType,
)

# Support both package and standalone imports
try:
    from ..function_type import FunctionTypeAnalysis
    from ..utils.hashes import calculate_sha256
except ImportError:
    # Fallback to absolute imports (for multiprocessing spawned processes)
    from redb.extractors.decompiler.bninja.function_type import FunctionTypeAnalysis
    from redb.extractors.decompiler.bninja.utils.hashes import calculate_sha256


class DisassemblyAnalysis:
    INVALID_STACK_SIZE = -1

    def __init__(self, arch, function, bv, logger):
        self.arch = arch
        self.function = function
        self.bv = bv
        self.logger = logger
        if self.function is not None and hasattr(self.function, "instructions"):
            self.instructions = self.function.instructions
        else:
            self.instructions = []
        self.errors = []
        return

    def log_error(
        self, message, function_name, address, exception=None, error_location="unknown"
    ):
        """Log an error during processing."""
        error_msg = f"Error in function {function_name} at {address}: {message}"
        if exception:
            error_msg += f" - {str(exception)}"
        self.logger.error(error_msg)

        # Add to errors list
        error = {
            "function_name": function_name,
            "function_address": str(address),
            "error_location": error_location,
            "error_message": message,
            "error_details": str(exception) if exception else "",
            "error_type": type(exception).__name__ if exception else "Unknown",
            "timestamp": int(time.time() * 1000),
        }
        self.errors.append(error)

    def get_json(self):
        try:
            # Build disassembly string and normalized versions
            disassembly_builder = [[], []]  # Address and instruction text

            # Create a dictionary mapping addresses to instruction tokens
            instr_tokens_by_addr = {}
            for instr_tokens, addr in self.instructions:
                instr_tokens_by_addr[addr] = instr_tokens

            addresses = sorted(instr_tokens_by_addr.keys())
            for address in addresses:
                # Original disassembly with addresses
                # instr_tokens, address = instruction
                instr_tokens = instr_tokens_by_addr[address]
                disassembly_builder[0].append(address)
                disassembly_builder[1].append("".join(map(str, instr_tokens)))

            # Join with newlines
            disassembly_str = "\n".join(disassembly_builder[1])
            disassembly_with_addresses = "\n".join(
                f"{hex(address)}: {instr_text}"
                for address, instr_text in zip(
                    disassembly_builder[0], disassembly_builder[1], strict=False
                )
            )

            disassembly_json = {
                "disassembled_function_hash": calculate_sha256(disassembly_str),
                "disassembled_function": disassembly_with_addresses,
                "disassembled_function_no_addresses": disassembly_str,
                "disassembled_function_name": self.function.name,
                "disassembled_function_address": self.function.start,
                "instructions_count": len(instr_tokens_by_addr.keys()),
                "function_type": FunctionTypeAnalysis(self.function)
                .get_function_type()
                .name,
            }

            # Add additional metrics
            type_frequencies = self.collect_instruction_types()
            disassembly_json["instructions_types"] = list(type_frequencies.keys())
            disassembly_json["control_flow_count"] = (
                self.count_control_flow_instructions()
            )
            disassembly_json["memory_access_pattern"] = self.collect_memory_patterns()
            disassembly_json["register_usage"] = self.collect_register_usage()
            disassembly_json["data_references_count"] = self.count_data_references()
            disassembly_json["max_block_size"] = self.compute_max_block_size()
            disassembly_json["num_calls"] = self.compute_num_calls()
            disassembly_json["stack_size"] = self.estimate_stack_size()

            return disassembly_json, self.errors

        except Exception as e:
            self.log_error(
                "Failed to collect instruction types",
                self.function.name,
                self.function.start,
                e,
                "collect_instruction_types",
            )
            raise ValueError(e) from e

    def collect_instruction_types(self):
        """Collect instruction type frequencies from a function."""
        type_frequencies = {}

        try:
            # Iterate through all instructions in the function
            for instruction in self.instructions:
                instr_tokens = instruction[0]  # Get the instruction tokens

                # Extract the mnemonic from the instruction tokens
                mnemonic = None
                for token in instr_tokens:
                    if token.type == InstructionTextTokenType.InstructionToken:
                        mnemonic = token.text
                        break

                if not mnemonic:
                    continue

                # Use normalize_opcode to get standardized opcode
                normalized = self.normalize_opcode(mnemonic)

                # Get category from opcode_categories or use the instruction type directly
                category = self.arch.opcode_categories.get(normalized)
                if category:
                    self._increment_frequency(type_frequencies, category)

        except Exception as e:
            self.log_error(
                "Failed to collect instruction types",
                self.function.name,
                self.function.start,
                e,
                "collect_instruction_types",
            )

        return type_frequencies

    def normalize_opcode(self, opcode):
        return opcode.upper()

    def collect_memory_patterns(self):
        """Collect memory access patterns from a function."""
        patterns = []
        try:
            for instruction in self.instructions:
                instr_tokens = instruction[0]

                # We need to capture memory operands between BeginMemoryOperandToken and EndMemoryOperandToken
                in_memory_operand = False
                memory_operand_text = ""

                for token in instr_tokens:
                    if token.type == InstructionTextTokenType.BeginMemoryOperandToken:
                        in_memory_operand = True
                        memory_operand_text = ""
                    elif token.type == InstructionTextTokenType.EndMemoryOperandToken:
                        in_memory_operand = False

                        # Process the captured memory operand text
                        if memory_operand_text:
                            # Categorize memory access pattern
                            if (
                                "+" in memory_operand_text
                                and "*" in memory_operand_text
                            ):
                                if "MEM_SCALED_INDEX" not in patterns:
                                    patterns.append("MEM_SCALED_INDEX")
                            elif (
                                "+" in memory_operand_text or "-" in memory_operand_text
                            ):
                                if "MEM_BASE_OFFSET" not in patterns:
                                    patterns.append("MEM_BASE_OFFSET")
                            else:
                                if "MEM_DIRECT" not in patterns:
                                    patterns.append("MEM_DIRECT")

                            # Check for stack accesses
                            if any(
                                reg in memory_operand_text
                                for reg in ["SP", "BP", "ESP", "EBP", "RSP", "RBP"]
                            ):
                                if "MEM_STACK" not in patterns:
                                    patterns.append("MEM_STACK")
                            # Check for string operations
                            elif (
                                any(
                                    reg in memory_operand_text
                                    for reg in ["SI", "DI", "ESI", "EDI", "RSI", "RDI"]
                                )
                                and "MEM_STRING" not in patterns
                            ):
                                patterns.append("MEM_STRING")
                    elif in_memory_operand:
                        # Accumulate token text while inside a memory operand
                        memory_operand_text += token.text
        except Exception as e:
            self.log_error(
                "Failed to collect memory patterns",
                self.function.name,
                self.function.start,
                e,
                "collect_memory_patterns",
            )
        return patterns

    def collect_register_usage(self):
        """Collect register usage from a function."""
        registers = []
        try:
            # Define register groups we're interested in tracking
            register_groups = {
                "GPR": [
                    "RAX",
                    "RBX",
                    "RCX",
                    "RDX",
                    "R9",
                    "R10",
                    "R11",
                    "R12",
                    "R13",
                    "R14",
                    "R15",
                    "EAX",
                    "EBX",
                    "ECX",
                    "EDX",
                    "R9D",
                    "R10D",
                    "R11D",
                    "R12D",
                    "R13D",
                    "R14D",
                    "AX",
                    "BX",
                    "CX",
                    "DX",
                ],
                "GPR_INDEX": ["RSI", "RDI", "ESI", "EDI", "SI", "DI"],
                "GPR_STACK": ["RSP", "RBP", "ESP", "EBP", "SP", "BP"],
                "SIMD": ["XMM", "YMM", "ZMM"],
                "FPU": ["ST", "ST0", "ST1", "ST2", "ST3", "ST4", "ST5", "ST6", "ST7"],
                "FLAGS": ["FLAGS", "EFLAGS", "RFLAGS"],
                "CONTROL_REGISTER": ["CR0", "CR2", "CR3", "CR4", "CR8"],
                "DEBUG_REGISTER": ["DR0", "DR1", "DR2", "DR3", "DR6", "DR7"],
            }

            # Extract registers from instructions
            for instruction in self.instructions:
                instr_tokens = instruction[0]
                for token in instr_tokens:
                    if token.type == InstructionTextTokenType.RegisterToken:
                        reg = token.text.upper()
                        # Check which group this register belongs to
                        for group, regs in register_groups.items():
                            # if any(r in reg for r in regs) or any(reg.startswith(r) for r in regs):
                            if any(reg == r or reg.startswith(r) for r in regs):
                                if group not in registers:
                                    registers.append(group)
                                break
        except Exception as e:
            self.log_error(
                "Failed to collect register usage",
                self.function.name,
                self.function.start,
                e,
                "collect_register_usage",
            )
        return registers

    def count_data_references(self):
        """Count the number of data references in a function."""
        count = 0
        try:

            if self.function.mlil is None:
                return 0

            for block in self.function.mlil:
                for instr in block:
                    instr_str = str(instr)
                    logged = False
                    src = None

                    # Check for constant dereferencing or symbolic refs
                    if hasattr(instr, "src"):
                        src = instr.src
                        if isinstance(
                            src,
                            (
                                binaryninja.mediumlevelil.MediumLevelILConstPtr,
                                binaryninja.mediumlevelil.MediumLevelILConst,
                            ),
                        ):
                            count += 1
                            logged = True

                    # Check full string for hardcoded addresses or symbol-like tokens
                    if re.search(r"\b0x[0-9A-Fa-f]{3,}\b", instr_str) and not logged:
                        count += 1
                        logged = True

                    if "_" in instr_str and not logged:
                        count += 1
                        logged = True

                    # Only check for MediumLevelILConstPtr if src exists
                    if src is not None and isinstance(
                        src, binaryninja.mediumlevelil.MediumLevelILConstPtr
                    ):
                        addr = src.constant
                        # Check if address is in data sections
                        segment = self.bv.get_segment_at(addr)
                        if segment and segment.writable:
                            # print(f"[{function.name}] Matched data section reference in: {instr_str}")
                            count += 1
                            logged = True
        except Exception as e:
            self.logger.warning(
                f"Failed to use MLIL for counting data references in {self.function.name} at {self.function.start}: {e}"
            )
        return count

    def compute_max_block_size(self):
        """Compute the maximum basic block size in a function."""
        max_size = 0
        if self.function is None:
            return 0

        for block in self.function.basic_blocks:
            try:
                # Count instructions in this block using the direct length approach
                # This avoids UTF-8 decoding issues entirely
                block_size = block.instruction_count
                max_size = max(max_size, block_size)
            except Exception as e:
                self.log_error(
                    f"[HandledError] computing max block size: {e}",
                    self.function.name,
                    self.function.start,
                    e,
                    "compute_max_block_size",
                )
        return max_size

    def count_control_flow_instructions(self):
        """Count the number of control flow instructions in a function."""
        count = 0
        try:
            for instruction in self.instructions:
                instr_tokens = instruction[0]
                if self.arch.is_control_flow_instruction(instr_tokens):
                    count += 1
        except Exception as e:
            self.log_error(
                "Failed to count control flow instructions",
                self.function.name,
                self.function.start,
                e,
                "count_control_flow_instructions",
            )
        return count

    def compute_num_calls(self) -> int:
        """Compute the number of call instructions in a function."""
        num_calls = 0
        try:
            for instruction in self.instructions:
                instr_tokens = instruction[0]
                # Extract the mnemonic
                for token in instr_tokens:
                    if token.type == InstructionTextTokenType.InstructionToken:
                        if token.text.upper() == "CALL":
                            num_calls += 1
                        break
        except Exception as e:
            self.log_error(
                "Failed to compute number of calls",
                self.function.name,
                self.function.start,
                e,
                "compute_num_calls",
            )
        return num_calls

    def _increment_frequency(self, frequencies, type_name):
        """Increment the frequency count for an instruction type."""
        if type_name in frequencies:
            frequencies[type_name] += 1
        else:
            frequencies[type_name] = 1

    def estimate_stack_size(self):
        """Estimate the stack size used by a function."""
        try:
            # Binary Ninja provides a stack adjustment value for functions
            # Need to convert OffsetWithConfidence to a plain integer
            stack_adjust = self.function.stack_adjustment
            if hasattr(stack_adjust, "value"):  # Handle OffsetWithConfidence objects
                return stack_adjust.value
            return stack_adjust
        except Exception as e:
            self.log_error(
                "Failed to estimate stack size",
                self.function.name,
                self.function.start,
                e,
                "estimate_stack_size",
            )
            return self.INVALID_STACK_SIZE