Wei Kong

73 papers A* 3A 1B 2C 6Journal 42Unranked 19
YearRankTypeTitle / Venue / Authors
2025 J jnl
Symmetry
Han Wang, Fengxiang Wang, Ruikai Xue, Xiaokai She, Wei Kong, Genghua Huang
2025 J jnl
IEEE Trans. Geosci. Remote. Sens.
Su Chen, Peng Chen, Wei Kong, Rong Shu, Delu Pan
2025 A* conf
NDSS
Jianwen Tian, Wei Kong, Debin Gao, Tong Wang, Taotao Gu, Kefan Qiu, Zhi Wang, Xiaohui Kuang
2025 J jnl
J. Comput. Biol.
Shunqin Zhang, Wei Kong, Shuaiqun Wang, Kai Wei, Kun Liu, Gen Wen, Yaling Yu
2025 conf
CSS
Jiahe Ji, Wei Kong, Yang Liu
2025 J jnl
Array
Xiaoyu Chen, Shuaiqun Wang, Wei Kong
2025 J jnl
IEEE Trans. Geosci. Remote. Sens.
Su Chen, Peng Chen, Wei Kong, Rong Shu, Delu Pan
2025 J jnl
Neurocomputing
Jin Deng, Kaihan Huang, Jinfeng Wang, Wenjian Zhong, Yufang Xu, Wei Kong
2025 conf
BIBM
Jin Deng, Tao Xu, Jianjun Zhang, Lechun Liu, Wei Kong
2024 A* conf
ICRA
Wei Kong, Hu Li, Qianjin Du, Huayang Cao, Xiaohui Kuang
2024 J jnl
IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens.
Kuifeng Luan, Xueyan Zhao, Wei Kong, Tao Chen, Huan Xie, Xiangfeng Liu, Fengxiang Wang
2024 J jnl
IEEE Trans. Geosci. Remote. Sens.
Danchen Wu, Peng Chen, Wei Kong, Delu Pan
2024 B conf
TrustCom
Yongkang Chen, Tong Wang, Wei Kong, Taotao Gu, Guiling Cao, Xiaohui Kuang
2024 conf
BIC
Yawen Chen, Wei Kong
2024 J jnl
Appl. Intell.
Xiaoqian Hu, Yaling Yu, Wei Kong, Shuaiqun Wang, Gen Wen
2024 conf
ML4CS
Wei Wu, Chen Wang, Qiuhao Xu, Wei Kong
2024 A* conf
NeurIPS
Jingjing Wang, Minhuan Huang, Yuanping Nie, Xiang Li, Qianjin Du, Wei Kong, Huan Deng, Xiaohui Kuang
2023 conf
ISAIMS
Xiaoqian Hu, Gen Wen, Wei Kong
2023 J jnl
Genom. Proteom. Bioinform.
Yicong Shen, Yuanxu Gao, Jiangcheng Shi, Zhou Huang, Rongbo Dai, Yi Fu, Yuan Zhou, Wei Kong, Qinghua Cui
2023 J jnl
Briefings Bioinform.
Shuaiqun Wang, Kai Zheng, Wei Kong, Ruiwen Huang, Lulu Liu, Gen Wen, Yaling Yu
2023 B conf
IJCNN
Tong Wang, Xiaohui Kuang, Huan Deng, Taotao Gu, Wei Kong, Jianwen Tian, Gang Zhao
2023 conf
DSC
Jiahe Ji, Wei Kong, Jianwen Tian, Taotao Gu, Yuanping Nie, Xiaohui Kuang
2023 C conf
APNOMS
Wei Kong, Qianjin Du, Huayang Cao, Hu Li, Tong Wang, Jianwen Tian, Xiaohui Kuang
2023 C conf
ICCD
Wei Kong
2022 J jnl
Remote. Sens.
Xue Shen, Wei Kong, Peng Chen, Tao Chen, Genghua Huang, Rong Shu
2022 J jnl
ACM Comput. Surv.
Huawei Huang, Wei Kong, Sicong Zhou, Zibin Zheng, Song Guo
2022 conf
DSC
Hanwen Sun, Yuanping Nie, Xiang Li, Minhuan Huang, Jianwen Tian, Wei Kong
2022 J jnl
Medical Biol. Eng. Comput.
Kai Wei, Wei Kong, Shuaiqun Wang
2022 J jnl
Remote. Sens.
Changsheng Tan, Wei Kong, Genghua Huang, Jia Hou, Shaolei Jia, Tao Chen, Rong Shu
2022 J jnl
Remote. Sens.
Wenjie Yue, Tao Chen, Wei Kong, Xin Chen, Genghua Huang, Rong Shu
2022 J jnl
Pattern Recognit. Lett.
Wei Kong, Yun Liu, Hui Li, Chuanxu Wang, Ye Tao, Xiangzhen Kong
2022 J jnl
Remote. Sens.
Rujia Ma, Wei Kong, Tao Chen, Rong Shu, Genghua Huang
2022 J jnl
Sensors
Yun Xiao, Wei Kong, Zijun Liang
2021 J jnl
IEEE Access
Kai Wei, Wei Kong, Shuaiqun Wang
2021 conf
DSC
Wei Kong, Huayang Cao, Jianwen Tian, Xiaohui Kuang
2021 J jnl
J. Bioinform. Comput. Biol.
Lei Wang, Wei Kong, Shuaiqun Wang
2021 J jnl
Comput. Intell. Neurosci.
Wei Kong, Yun Liu, Hui Li, Chuanxu Wang
2021 J jnl
Comput. Math. Methods Medicine
Shuaiqun Wang, Xiaoling Xu, Wei Kong
2021 J jnl
Inf. Sci.
Jin Deng, Weiming Zeng, Sizhe Luo, Wei Kong, Yuhu Shi, Ying Li, Hua Zhang
2021 conf
ICCBN
Zijun Liang, Haoyun Liu, Wei Kong, Ting Lv
2020 J jnl
CoRR
Huawei Huang, Wei Kong, Sicong Zhou, Zibin Zheng, Song Guo
2020 J jnl
J. Parallel Distributed Comput.
Wei Kong, Jian Shen, Pandi Vijayakumar, Youngju Cho, Victor Chang
2020 J jnl
Comput. Math. Methods Medicine
Jin Deng, Weiming Zeng, Yuhu Shi, Wei Kong, Shunjie Guo
2020 J jnl
IEEE Trans. Biomed. Eng.
Jin Deng, Weiming Zeng, Wei Kong, Yuhu Shi, Xiaoyang Mou, Jian Guo
2019 J jnl
IEEE Access
Yuling Chen, Wei Kong, Xinzhao Jiang
2018 J jnl
Medical Biol. Eng. Comput.
Yuhu Shi, Weiming Zeng, Xiaoyan Tang, Wei Kong, Jun Yin
2017 J jnl
BMC Genom.
Yang Yang, Ning Huang, Luning Hao, Wei Kong
2017 J jnl
Briefings Bioinform.
Wei Ma, Lu Zhang, Pan Zeng, Chuanbo Huang, Jianwei Li, Bin Geng, Jichun Yang, Wei Kong, Xuezhong Zhou, Qinghua Cui
2017 J jnl
Int. J. Data Min. Bioinform.
Yang Yang, Yiqun Xiao, Tianyu Cao, Wei Kong
2016 conf
BIBM
Yang Yang, Tianyu Cao, Wei Kong
2014 J jnl
Comput. Math. Methods Medicine
Wei Kong, Xiaoyang Mou, Xing Zhi, Xing Zhang, Yang Yang
2014 J jnl
Comput. Math. Methods Medicine
Wei Kong, Jingmao Zhang, Xiaoyang Mou, Yang Yang
2013 conf
ICCI*CC
Ling Nie, Shali Xiao, Miao Li, Xi Wang, Junling Zhang, Wei Kong
2012 J jnl
Simul.
Wei Kong, Da-qiang Cang
2011 conf
A-SSCC
Yasunobu Nakase, Shinichi Hirose, Toru Goda, Kehui Hu, Hiroshi Onoda, Yasuhiro Ido, Hiroyuki Kondo, Wei Kong, Wei Zhang, Tsukasa Oishi, Shintaro Mori, Toru Shimizu
2011 J jnl
IEEE Trans. Biomed. Eng.
Wei Kong, Andrew E. Pollard, Vladimir G. Fast
2011 conf
HCI (8)
Wei Kong, Josette F. Jones
2011 J jnl
BMC Bioinform.
Wei Kong, Xiaoyang Mou, Xiaohua Hu
2011 conf
iCAST
Wei Kong, Xiaoyang Mou
2010 conf
BIBM
Wei Kong, Xiaoyang Mou, Xiaohua Hu
2010 J jnl
Artif. Intell. Medicine
Qi Shen, Wei-Min Shi, Wei Kong
2009 conf
IJCBS
Hui Liu, Wei Kong, Tianshuang Qiu, Guo-li Li
2009 J jnl
J. Biomed. Informatics
Qi Shen, Wei-Min Shi, Wei Kong
2009 conf
IJCBS
Wei Kong, Xiaoyang Mou, Bin Yang
2008 A conf
ITC
Wei Kong, Paul C. Parries, G. Wang, Subramanian S. Iyer
2008 J jnl
Comput. Biol. Chem.
Qi Shen, Wei-Min Shi, Wei Kong
2007 C conf
IPCCC
Xiaolin Chang, Jogesh K. Muppala, Wei Kong, Pengcheng Zou, Xiangkai Li, Zhongyuan Zheng
2006 conf
ISNN (2)
Wei Kong, Bin Yang
2005 C conf
CIARP
Wei Kong, Yang Bin
2005 J jnl
IEEE Trans. Biomed. Eng.
Wei Kong, Dennis L. Rollins, Raymond E. Ideker, William M. Smith
2005 C conf
CIARP
Yang Bin, Wei Kong
2004 conf
Australian Conference on Artificial Intelligence
Yue Zhou, Wei Kong, Qing Xu
2004 C conf
CIARP
Wei Kong, Yue Zhou, Jie Yang
redb/extractors/decompiler/bninja/analysis/disassembly.py
← Index redb/extractors/decompiler/bninja/analysis/disassembly.py python
import re
import time

import binaryninja
from binaryninja.enums import (
    InstructionTextTokenType,
)

# Support both package and standalone imports
try:
    from ..function_type import FunctionTypeAnalysis
    from ..utils.hashes import calculate_sha256
except ImportError:
    # Fallback to absolute imports (for multiprocessing spawned processes)
    from redb.extractors.decompiler.bninja.function_type import FunctionTypeAnalysis
    from redb.extractors.decompiler.bninja.utils.hashes import calculate_sha256


class DisassemblyAnalysis:
    INVALID_STACK_SIZE = -1

    def __init__(self, arch, function, bv, logger):
        self.arch = arch
        self.function = function
        self.bv = bv
        self.logger = logger
        if self.function is not None and hasattr(self.function, "instructions"):
            self.instructions = self.function.instructions
        else:
            self.instructions = []
        self.errors = []
        return

    def log_error(
        self, message, function_name, address, exception=None, error_location="unknown"
    ):
        """Log an error during processing."""
        error_msg = f"Error in function {function_name} at {address}: {message}"
        if exception:
            error_msg += f" - {str(exception)}"
        self.logger.error(error_msg)

        # Add to errors list
        error = {
            "function_name": function_name,
            "function_address": str(address),
            "error_location": error_location,
            "error_message": message,
            "error_details": str(exception) if exception else "",
            "error_type": type(exception).__name__ if exception else "Unknown",
            "timestamp": int(time.time() * 1000),
        }
        self.errors.append(error)

    def get_json(self):
        try:
            # Build disassembly string and normalized versions
            disassembly_builder = [[], []]  # Address and instruction text

            # Create a dictionary mapping addresses to instruction tokens
            instr_tokens_by_addr = {}
            for instr_tokens, addr in self.instructions:
                instr_tokens_by_addr[addr] = instr_tokens

            addresses = sorted(instr_tokens_by_addr.keys())
            for address in addresses:
                # Original disassembly with addresses
                # instr_tokens, address = instruction
                instr_tokens = instr_tokens_by_addr[address]
                disassembly_builder[0].append(address)
                disassembly_builder[1].append("".join(map(str, instr_tokens)))

            # Join with newlines
            disassembly_str = "\n".join(disassembly_builder[1])
            disassembly_with_addresses = "\n".join(
                f"{hex(address)}: {instr_text}"
                for address, instr_text in zip(
                    disassembly_builder[0], disassembly_builder[1], strict=False
                )
            )

            disassembly_json = {
                "disassembled_function_hash": calculate_sha256(disassembly_str),
                "disassembled_function": disassembly_with_addresses,
                "disassembled_function_no_addresses": disassembly_str,
                "disassembled_function_name": self.function.name,
                "disassembled_function_address": self.function.start,
                "instructions_count": len(instr_tokens_by_addr.keys()),
                "function_type": FunctionTypeAnalysis(self.function)
                .get_function_type()
                .name,
            }

            # Add additional metrics
            type_frequencies = self.collect_instruction_types()
            disassembly_json["instructions_types"] = list(type_frequencies.keys())
            disassembly_json["control_flow_count"] = (
                self.count_control_flow_instructions()
            )
            disassembly_json["memory_access_pattern"] = self.collect_memory_patterns()
            disassembly_json["register_usage"] = self.collect_register_usage()
            disassembly_json["data_references_count"] = self.count_data_references()
            disassembly_json["max_block_size"] = self.compute_max_block_size()
            disassembly_json["num_calls"] = self.compute_num_calls()
            disassembly_json["stack_size"] = self.estimate_stack_size()

            return disassembly_json, self.errors

        except Exception as e:
            self.log_error(
                "Failed to collect instruction types",
                self.function.name,
                self.function.start,
                e,
                "collect_instruction_types",
            )
            raise ValueError(e) from e

    def collect_instruction_types(self):
        """Collect instruction type frequencies from a function."""
        type_frequencies = {}

        try:
            # Iterate through all instructions in the function
            for instruction in self.instructions:
                instr_tokens = instruction[0]  # Get the instruction tokens

                # Extract the mnemonic from the instruction tokens
                mnemonic = None
                for token in instr_tokens:
                    if token.type == InstructionTextTokenType.InstructionToken:
                        mnemonic = token.text
                        break

                if not mnemonic:
                    continue

                # Use normalize_opcode to get standardized opcode
                normalized = self.normalize_opcode(mnemonic)

                # Get category from opcode_categories or use the instruction type directly
                category = self.arch.opcode_categories.get(normalized)
                if category:
                    self._increment_frequency(type_frequencies, category)

        except Exception as e:
            self.log_error(
                "Failed to collect instruction types",
                self.function.name,
                self.function.start,
                e,
                "collect_instruction_types",
            )

        return type_frequencies

    def normalize_opcode(self, opcode):
        return opcode.upper()

    def collect_memory_patterns(self):
        """Collect memory access patterns from a function."""
        patterns = []
        try:
            for instruction in self.instructions:
                instr_tokens = instruction[0]

                # We need to capture memory operands between BeginMemoryOperandToken and EndMemoryOperandToken
                in_memory_operand = False
                memory_operand_text = ""

                for token in instr_tokens:
                    if token.type == InstructionTextTokenType.BeginMemoryOperandToken:
                        in_memory_operand = True
                        memory_operand_text = ""
                    elif token.type == InstructionTextTokenType.EndMemoryOperandToken:
                        in_memory_operand = False

                        # Process the captured memory operand text
                        if memory_operand_text:
                            # Categorize memory access pattern
                            if (
                                "+" in memory_operand_text
                                and "*" in memory_operand_text
                            ):
                                if "MEM_SCALED_INDEX" not in patterns:
                                    patterns.append("MEM_SCALED_INDEX")
                            elif (
                                "+" in memory_operand_text or "-" in memory_operand_text
                            ):
                                if "MEM_BASE_OFFSET" not in patterns:
                                    patterns.append("MEM_BASE_OFFSET")
                            else:
                                if "MEM_DIRECT" not in patterns:
                                    patterns.append("MEM_DIRECT")

                            # Check for stack accesses
                            if any(
                                reg in memory_operand_text
                                for reg in ["SP", "BP", "ESP", "EBP", "RSP", "RBP"]
                            ):
                                if "MEM_STACK" not in patterns:
                                    patterns.append("MEM_STACK")
                            # Check for string operations
                            elif (
                                any(
                                    reg in memory_operand_text
                                    for reg in ["SI", "DI", "ESI", "EDI", "RSI", "RDI"]
                                )
                                and "MEM_STRING" not in patterns
                            ):
                                patterns.append("MEM_STRING")
                    elif in_memory_operand:
                        # Accumulate token text while inside a memory operand
                        memory_operand_text += token.text
        except Exception as e:
            self.log_error(
                "Failed to collect memory patterns",
                self.function.name,
                self.function.start,
                e,
                "collect_memory_patterns",
            )
        return patterns

    def collect_register_usage(self):
        """Collect register usage from a function."""
        registers = []
        try:
            # Define register groups we're interested in tracking
            register_groups = {
                "GPR": [
                    "RAX",
                    "RBX",
                    "RCX",
                    "RDX",
                    "R9",
                    "R10",
                    "R11",
                    "R12",
                    "R13",
                    "R14",
                    "R15",
                    "EAX",
                    "EBX",
                    "ECX",
                    "EDX",
                    "R9D",
                    "R10D",
                    "R11D",
                    "R12D",
                    "R13D",
                    "R14D",
                    "AX",
                    "BX",
                    "CX",
                    "DX",
                ],
                "GPR_INDEX": ["RSI", "RDI", "ESI", "EDI", "SI", "DI"],
                "GPR_STACK": ["RSP", "RBP", "ESP", "EBP", "SP", "BP"],
                "SIMD": ["XMM", "YMM", "ZMM"],
                "FPU": ["ST", "ST0", "ST1", "ST2", "ST3", "ST4", "ST5", "ST6", "ST7"],
                "FLAGS": ["FLAGS", "EFLAGS", "RFLAGS"],
                "CONTROL_REGISTER": ["CR0", "CR2", "CR3", "CR4", "CR8"],
                "DEBUG_REGISTER": ["DR0", "DR1", "DR2", "DR3", "DR6", "DR7"],
            }

            # Extract registers from instructions
            for instruction in self.instructions:
                instr_tokens = instruction[0]
                for token in instr_tokens:
                    if token.type == InstructionTextTokenType.RegisterToken:
                        reg = token.text.upper()
                        # Check which group this register belongs to
                        for group, regs in register_groups.items():
                            # if any(r in reg for r in regs) or any(reg.startswith(r) for r in regs):
                            if any(reg == r or reg.startswith(r) for r in regs):
                                if group not in registers:
                                    registers.append(group)
                                break
        except Exception as e:
            self.log_error(
                "Failed to collect register usage",
                self.function.name,
                self.function.start,
                e,
                "collect_register_usage",
            )
        return registers

    def count_data_references(self):
        """Count the number of data references in a function."""
        count = 0
        try:

            if self.function.mlil is None:
                return 0

            for block in self.function.mlil:
                for instr in block:
                    instr_str = str(instr)
                    logged = False
                    src = None

                    # Check for constant dereferencing or symbolic refs
                    if hasattr(instr, "src"):
                        src = instr.src
                        if isinstance(
                            src,
                            (
                                binaryninja.mediumlevelil.MediumLevelILConstPtr,
                                binaryninja.mediumlevelil.MediumLevelILConst,
                            ),
                        ):
                            count += 1
                            logged = True

                    # Check full string for hardcoded addresses or symbol-like tokens
                    if re.search(r"\b0x[0-9A-Fa-f]{3,}\b", instr_str) and not logged:
                        count += 1
                        logged = True

                    if "_" in instr_str and not logged:
                        count += 1
                        logged = True

                    # Only check for MediumLevelILConstPtr if src exists
                    if src is not None and isinstance(
                        src, binaryninja.mediumlevelil.MediumLevelILConstPtr
                    ):
                        addr = src.constant
                        # Check if address is in data sections
                        segment = self.bv.get_segment_at(addr)
                        if segment and segment.writable:
                            # print(f"[{function.name}] Matched data section reference in: {instr_str}")
                            count += 1
                            logged = True
        except Exception as e:
            self.logger.warning(
                f"Failed to use MLIL for counting data references in {self.function.name} at {self.function.start}: {e}"
            )
        return count

    def compute_max_block_size(self):
        """Compute the maximum basic block size in a function."""
        max_size = 0
        if self.function is None:
            return 0

        for block in self.function.basic_blocks:
            try:
                # Count instructions in this block using the direct length approach
                # This avoids UTF-8 decoding issues entirely
                block_size = block.instruction_count
                max_size = max(max_size, block_size)
            except Exception as e:
                self.log_error(
                    f"[HandledError] computing max block size: {e}",
                    self.function.name,
                    self.function.start,
                    e,
                    "compute_max_block_size",
                )
        return max_size

    def count_control_flow_instructions(self):
        """Count the number of control flow instructions in a function."""
        count = 0
        try:
            for instruction in self.instructions:
                instr_tokens = instruction[0]
                if self.arch.is_control_flow_instruction(instr_tokens):
                    count += 1
        except Exception as e:
            self.log_error(
                "Failed to count control flow instructions",
                self.function.name,
                self.function.start,
                e,
                "count_control_flow_instructions",
            )
        return count

    def compute_num_calls(self) -> int:
        """Compute the number of call instructions in a function."""
        num_calls = 0
        try:
            for instruction in self.instructions:
                instr_tokens = instruction[0]
                # Extract the mnemonic
                for token in instr_tokens:
                    if token.type == InstructionTextTokenType.InstructionToken:
                        if token.text.upper() == "CALL":
                            num_calls += 1
                        break
        except Exception as e:
            self.log_error(
                "Failed to compute number of calls",
                self.function.name,
                self.function.start,
                e,
                "compute_num_calls",
            )
        return num_calls

    def _increment_frequency(self, frequencies, type_name):
        """Increment the frequency count for an instruction type."""
        if type_name in frequencies:
            frequencies[type_name] += 1
        else:
            frequencies[type_name] = 1

    def estimate_stack_size(self):
        """Estimate the stack size used by a function."""
        try:
            # Binary Ninja provides a stack adjustment value for functions
            # Need to convert OffsetWithConfidence to a plain integer
            stack_adjust = self.function.stack_adjustment
            if hasattr(stack_adjust, "value"):  # Handle OffsetWithConfidence objects
                return stack_adjust.value
            return stack_adjust
        except Exception as e:
            self.log_error(
                "Failed to estimate stack size",
                self.function.name,
                self.function.start,
                e,
                "estimate_stack_size",
            )
            return self.INVALID_STACK_SIZE