Raffaele Modugno

13 papers A 3B 4Journal 1Unranked 4
YearRankTypeTitle / Venue / Authors
2012 B conf
ICFHR
Donato Impedovo, Giuseppe Pirlo, Raffaele Modugno
2010 B conf
ICFHR
Sebastiano Impedovo, Raffaele Modugno, Giuseppe Pirlo
2010 B conf
ICPR
Sebastiano Impedovo, Raffaele Modugno, Giuseppe Pirlo
2010 B conf
ICFHR
Sebastiano Impedovo, Giuseppe Pirlo, Raffaele Modugno, Anna Ferrante
2009 conf
AI*IA
Sebastiano Impedovo, Anna Ferrante, Raffaele Modugno, Giuseppe Pirlo
2009 A conf
ICDAR
Sebastiano Impedovo, Anna Ferrante, Raffaele Modugno
2009 A conf
ICDAR
Sebastiano Impedovo, Raffaele Modugno, Anna Ferrante, Erasmo Stasolla
2004 conf
IWFHR
Giovanni Dimauro, Sebastiano Impedovo, M. G. Lucchese, Raffaele Modugno, Giuseppe Pirlo
2003 A conf
ICDAR
Giovanni Dimauro, Sebastiano Impedovo, Raffaele Modugno, Giuseppe Pirlo
2003 J jnl
IEEE Trans. Circuits Syst. II Express Briefs
Giovanni Dimauro, Sebastiano Impedovo, Raffaele Modugno, Giuseppe Pirlo, R. Stefanelli
2002 conf
IWFHR
Giovanni Dimauro, Sebastiano Impedovo, Raffaele Modugno, Giuseppe Pirlo
2002 conf
IWFHR
Giovanni Dimauro, Sebastiano Impedovo, Raffaele Modugno, Giuseppe Pirlo, L. Sarcinella
2002
Raffaele Modugno
redb/extractors/decompiler/bninja/analysis/strings.py
← Index redb/extractors/decompiler/bninja/analysis/strings.py python
from collections import Counter
import math

class StringAnalysis:
    def __init__(self, bv, functions):
        self.bv = bv
        self.functions = functions

    def entropy(self, s: str) -> float:
        """Compute Shannon entropy of a string."""
        if not s:
            return 0.0
        freq = Counter(s)
        length = len(s)
        return -sum((count / length) * math.log2(count / length) for count in freq.values())

    def analyze(self):
        """
        Extract unique strings from the binary.

        Deduplicates by (string, encoding) within the same binary, keeping the
        first occurrence (lowest offset). Cross-binary deduplication and
        aggregation is handled by ClickHouse materialized views.
        """
        strings = {}

        # Sort strings by their starting address
        sorted_entries = sorted(self.bv.strings, key=lambda e: e.start)

        for entry in sorted_entries:
            # Key is the string and its encoding
            key = (entry.value, entry.type.name)

            # Skip if this string (value + encoding) was already added.
            # Because entries are sorted by address, the first one is always kept.
            if key in strings:
                continue

            # Store only the first occurrence with schema-matching field names
            # entry.length is the raw byte length, len(entry.value) is decoded string length
            string_entry = {
                "string": entry.value,
                "string_raw": entry.raw,
                "string_encoding": entry.type.name,
                "string_offset": entry.start,
                "string_length": len(entry.value),
                "string_raw_length": entry.length,
                "string_entropy": self.entropy(entry.value),
            }

            strings[key] = string_entry

        # Return as list for export compatibility
        return list(strings.values())