Hamdi Ben Hamadou

12 papers A 1B 1C 1Journal 2Unranked 6
YearRankTypeTitle / Venue / Authors
2021 J jnl
VLDB J.
Chiara Forresi, Enrico Gallinucci, Matteo Golfarelli, Hamdi Ben Hamadou
2020 conf
IEEE BigData
Hamdi Ben Hamadou, Torben Bach Pedersen, Christian Thomsen
2019 A conf
ER
Hamdi Ben Hamadou, Enrico Gallinucci, Matteo Golfarelli
2019
Hamdi Ben Hamadou
2019 J jnl
Inf. Syst.
Hamdi Ben Hamadou, Faiza Ghozzi, André Péninou, Olivier Teste
2018 conf
INFORSID
Mohammed El Malki, Hamdi Ben Hamadou, Max Chevalier, André Péninou, Olivier Teste
2018 conf
EGC
Hamdi Ben Hamadou, Faiza Ghozzi, André Péninou, Olivier Teste
2018 conf
ICEIS (1)
Mohammed El Malki, Hamdi Ben Hamadou, Nabil El Malki, Arlind Kopliku
2018 B conf
DaWaK
Mohammed El Malki, Hamdi Ben Hamadou, Max Chevalier, André Péninou, Olivier Teste
2018 conf
ICEIS (1)
Hamdi Ben Hamadou, Faiza Ghozzi, André Péninou, Olivier Teste
2018 conf
ICEIS (Revised Selected Papers)
Hamdi Ben Hamadou, Faiza Ghozzi, André Péninou, Olivier Teste
2018 C conf
DOLAP
Hamdi Ben Hamadou, Faiza Ghozzi, André Péninou, Olivier Teste
redb/extractors/decompiler/bninja/analysis/strings.py
← Index redb/extractors/decompiler/bninja/analysis/strings.py python
from collections import Counter
import math

class StringAnalysis:
    def __init__(self, bv, functions):
        self.bv = bv
        self.functions = functions

    def entropy(self, s: str) -> float:
        """Compute Shannon entropy of a string."""
        if not s:
            return 0.0
        freq = Counter(s)
        length = len(s)
        return -sum((count / length) * math.log2(count / length) for count in freq.values())

    def analyze(self):
        """
        Extract unique strings from the binary.

        Deduplicates by (string, encoding) within the same binary, keeping the
        first occurrence (lowest offset). Cross-binary deduplication and
        aggregation is handled by ClickHouse materialized views.
        """
        strings = {}

        # Sort strings by their starting address
        sorted_entries = sorted(self.bv.strings, key=lambda e: e.start)

        for entry in sorted_entries:
            # Key is the string and its encoding
            key = (entry.value, entry.type.name)

            # Skip if this string (value + encoding) was already added.
            # Because entries are sorted by address, the first one is always kept.
            if key in strings:
                continue

            # Store only the first occurrence with schema-matching field names
            # entry.length is the raw byte length, len(entry.value) is decoded string length
            string_entry = {
                "string": entry.value,
                "string_raw": entry.raw,
                "string_encoding": entry.type.name,
                "string_offset": entry.start,
                "string_length": len(entry.value),
                "string_raw_length": entry.length,
                "string_entropy": self.entropy(entry.value),
            }

            strings[key] = string_entry

        # Return as list for export compatibility
        return list(strings.values())