Ida M. Pu

16 papers B 1C 3Misc 1Journal 5Unranked 6
YearRankTypeTitle / Venue / Authors
2023 J jnl
CoRR
Henry Musto, Daniel Stamate, Ida M. Pu, Daniel Stahl
2023 B conf
ICCCI
Henry Musto, Daniel Stamate, Ida M. Pu, Daniel Stahl
2023 J jnl
CoRR
Henry Musto, Daniel Stamate, Ida M. Pu, Daniel Stahl
2021 C conf
ICMLA
Henry Musto, Daniel Stamate, Ida M. Pu, Daniel Stahl
2021 Misc conf
AIAI
Rapheal Olaniyan, Daniel Stamate, Ida M. Pu
2021 C conf
EANN
Mihai Ermaliuc, Daniel Stamate, George D. Magoulas, Ida M. Pu
2019 conf
ICCCI (1)
Rapheal Olaniyan, Daniel Stamate, Ida M. Pu, Alexander Zamyatin, Anna Vashkel, Frédéric Maréchal
2018 conf
IPMU (3)
Daniel Stamate, Wajdi Alghamdi, Daniel Stahl, Ida M. Pu, Fionn Murtagh, Danielle Belgrave, Robin M. Murray, Marta Di Forti
2014 J jnl
J. Discrete Algorithms
Ida M. Pu, Daniel Stamate, Yuji Shen
2012 J jnl
NeuroImage
Yuji Shen, Yi-Ching Lynn Ho, Rishma Vidyasagar, George Balanos, Xavier Golay, Ida M. Pu, Risto A. Kauppinen
2012 conf
IPMU (3)
Daniel Stamate, Ida M. Pu
2011 C conf
WOWMOM
Ida M. Pu, Yuji Shen
2010 J jnl
Math. Comput. Sci.
Ida M. Pu, Yuji Shen
2010 conf
ICT
Ida M. Pu, Jinguk Kim, Yuji Shen
2009 conf
NTMS
Ida M. Pu, Yuji Shen
2008 conf
ADHOC-NOW
Ida M. Pu, Yuji Shen, Jinguk Kim
redb/extractors/decompiler/bninja/analysis/strings.py
← Index redb/extractors/decompiler/bninja/analysis/strings.py python
from collections import Counter
import math

class StringAnalysis:
    def __init__(self, bv, functions):
        self.bv = bv
        self.functions = functions

    def entropy(self, s: str) -> float:
        """Compute Shannon entropy of a string."""
        if not s:
            return 0.0
        freq = Counter(s)
        length = len(s)
        return -sum((count / length) * math.log2(count / length) for count in freq.values())

    def analyze(self):
        """
        Extract unique strings from the binary.

        Deduplicates by (string, encoding) within the same binary, keeping the
        first occurrence (lowest offset). Cross-binary deduplication and
        aggregation is handled by ClickHouse materialized views.
        """
        strings = {}

        # Sort strings by their starting address
        sorted_entries = sorted(self.bv.strings, key=lambda e: e.start)

        for entry in sorted_entries:
            # Key is the string and its encoding
            key = (entry.value, entry.type.name)

            # Skip if this string (value + encoding) was already added.
            # Because entries are sorted by address, the first one is always kept.
            if key in strings:
                continue

            # Store only the first occurrence with schema-matching field names
            # entry.length is the raw byte length, len(entry.value) is decoded string length
            string_entry = {
                "string": entry.value,
                "string_raw": entry.raw,
                "string_encoding": entry.type.name,
                "string_offset": entry.start,
                "string_length": len(entry.value),
                "string_raw_length": entry.length,
                "string_entropy": self.entropy(entry.value),
            }

            strings[key] = string_entry

        # Return as list for export compatibility
        return list(strings.values())