Immo O. Kerner

14 papers Journal 6Unranked 7
YearRankTypeTitle / Venue / Authors
2006 conf
Informatik in der DDR
Immo O. Kerner
2004 conf
Informatik in der DDR
Immo O. Kerner
1998 book
Wolfgang Bibel, Herbert Fiedler, Werner Grass, Peter Gorny, Günter Hotz, Immo O. Kerner, Rüdiger Reischuk, Friedrich Roithmayr
1997 J jnl
Inform. Spektrum
Hartmut Fritzsche, Immo O. Kerner
1994 conf
IFIP Congress (2)
Immo O. Kerner
1993 conf
Computer Mediated Education of Information Technology Professionals and Advanced End-Users
Immo O. Kerner
1993 conf
INFOS
Immo O. Kerner
1985 J jnl
J. Inf. Process. Cybern.
Immo O. Kerner
1980 J jnl
J. Inf. Process. Cybern.
Immo O. Kerner
1975 conf
Methods of Algorithmic Language Implementation
Immo O. Kerner
1975 J jnl
J. Inf. Process. Cybern.
Immo O. Kerner
1974 J jnl
J. Inf. Process. Cybern.
Immo O. Kerner, Hassan Al-Sheikh-Khalil
1966 J jnl
Commun. ACM
Immo O. Kerner
1962 conf
IFIP Congress
Immo O. Kerner
redb/extractors/decompiler/bninja/analysis/strings.py
← Index redb/extractors/decompiler/bninja/analysis/strings.py python
from collections import Counter
import math

class StringAnalysis:
    def __init__(self, bv, functions):
        self.bv = bv
        self.functions = functions

    def entropy(self, s: str) -> float:
        """Compute Shannon entropy of a string."""
        if not s:
            return 0.0
        freq = Counter(s)
        length = len(s)
        return -sum((count / length) * math.log2(count / length) for count in freq.values())

    def analyze(self):
        """
        Extract unique strings from the binary.

        Deduplicates by (string, encoding) within the same binary, keeping the
        first occurrence (lowest offset). Cross-binary deduplication and
        aggregation is handled by ClickHouse materialized views.
        """
        strings = {}

        # Sort strings by their starting address
        sorted_entries = sorted(self.bv.strings, key=lambda e: e.start)

        for entry in sorted_entries:
            # Key is the string and its encoding
            key = (entry.value, entry.type.name)

            # Skip if this string (value + encoding) was already added.
            # Because entries are sorted by address, the first one is always kept.
            if key in strings:
                continue

            # Store only the first occurrence with schema-matching field names
            # entry.length is the raw byte length, len(entry.value) is decoded string length
            string_entry = {
                "string": entry.value,
                "string_raw": entry.raw,
                "string_encoding": entry.type.name,
                "string_offset": entry.start,
                "string_length": len(entry.value),
                "string_raw_length": entry.length,
                "string_entropy": self.entropy(entry.value),
            }

            strings[key] = string_entry

        # Return as list for export compatibility
        return list(strings.values())