C. E. Veni Madhavan

59 papers A 1B 6C 1Misc 4Journal 23Unranked 22
YearRankTypeTitle / Venue / Authors
2019 J jnl
Multim. Tools Appl.
R. Shreelekshmi, M. Wilscy, C. E. Veni Madhavan
2017 J jnl
IACR Cryptol. ePrint Arch.
Srikanth Ch, C. E. Veni Madhavan, Kumar Swamy H. V.
2017 conf
PKIA
C. E. Veni Madhavan
2014 J jnl
J. Symb. Comput.
Srinivas Vivek, C. E. Veni Madhavan
2014 conf
KDIR
Prateek Nagwanshi, C. E. Veni Madhavan
2013 J jnl
Signal Image Video Process.
R. Shreelekshmi, M. Wilscy, C. E. Veni Madhavan
2013 conf
KDIR/KMIS
Siddhartha Banerjee, Nitin Kumar, C. E. Veni Madhavan
2012 J jnl
CoRR
Rama Badrinath, C. E. Veni Madhavan
2012 J jnl
Top. Cogn. Sci.
Sudarshan Iyengar, C. E. Veni Madhavan, Katharina A. Zweig, Abhiram Natarajan
2011 B conf
IDA
V. Suresh, Avanthi Krishnamurthy, Rama Badrinath, C. E. Veni Madhavan
2011 conf
WWW (Companion Volume)
V. Suresh, Ashok Veilumuthu, Avanthi Krishnamurthy, C. E. Veni Madhavan, Kaushik Nath, Sunil Arvindam
2011 A conf
ECIR
Rama Badrinath, Suresh Venkatasubramaniyan, C. E. Veni Madhavan
2010 conf
ICSAP
R. Shreelekshmi, M. Wilscy, C. E. Veni Madhavan
2010 conf
CNSA
R. Shreelekshmi, M. Wilscy, C. E. Veni Madhavan
2010 conf
PAKDD (1)
R. Arun, V. Suresh, C. E. Veni Madhavan, M. Narasimha Murty
2009 conf
ICUMT
V. Suresh, H. V. Shashidhara, C. E. Veni Madhavan
2009 conf
PReMI
R. Arun, V. Suresh, C. E. Veni Madhavan
2009 J jnl
CoRR
Sujit Gujar, C. E. Veni Madhavan
2009 conf
ICSC
R. Arun, V. Suresh, C. E. Veni Madhavan
2008 B conf
ACIVS
Anil Yekkala, C. E. Veni Madhavan
2008 B conf
ISIT
Gagan Garg, P. Vijay Kumar, C. E. Veni Madhavan
2008 C conf
SETA
Gagan Garg, P. Vijay Kumar, C. E. Veni Madhavan
2007 conf
PReMI
Anil Yekkala, C. E. Veni Madhavan
2007 B conf
Discovery Science
V. Suresh, Narayanan Raghupathy, B. Shekar, C. E. Veni Madhavan
2006 conf
MRCS
V. Suresh, S. Maria Sophia, C. E. Veni Madhavan
2005 J jnl
Int. J. Comput. Math.
Abhijit Das, C. E. Veni Madhavan
2005 Misc ed.
INDOCRYPT
Subhamoy Maitra, C. E. Veni Madhavan, Ramarathnam Venkatesan
2003 J jnl
Electron. Notes Discret. Math.
Indivar Gupta, Laxmi Narain, C. E. Veni Madhavan
2002 J jnl
Discret. Appl. Math.
P. Sreenivasa Kumar, C. E. Veni Madhavan
2002 J jnl
J. Algorithms
C. R. Subramanian, C. E. Veni Madhavan
2001 J jnl
Electron. Colloquium Comput. Complex.
N. S. Narayanaswamy, C. E. Veni Madhavan
2001 Misc conf
COCOON
N. S. Narayanaswamy, C. E. Veni Madhavan
2000 Misc conf
INDOCRYPT
R. Sai Anand, C. E. Veni Madhavan
1999 B conf
ISAAC
Abhijit Das, C. E. Veni Madhavan
1998 J jnl
Random Struct. Algorithms
C. R. Subramanian, Martin Fürer, C. E. Veni Madhavan
1998 J jnl
Discret. Appl. Math.
P. Sreenivasa Kumar, C. E. Veni Madhavan
1995 Misc conf
COCOON
S. Rengarajan, C. E. Veni Madhavan
1994 J jnl
Vis. Comput.
Subir Kumar Ghosh, Anil Maheshwari, Sudebkumar Prasant Pal, C. E. Veni Madhavan
1994 conf
FSTTCS
C. R. Subramanian, C. E. Veni Madhavan
1993 J jnl
Comput. Geom.
Subir Kumar Ghosh, Anil Maheshwari, Sudebkumar Prasant Pal, Sanjeev Saluja, C. E. Veni Madhavan
1993 B conf
ISAAC
Martin Fürer, C. R. Subramanian, C. E. Veni Madhavan
1993 conf
CCCG
P. Pradeep Kumar, C. E. Veni Madhavan
1992 J jnl
Inf. Process. Lett.
N. Ch. Veeraraghavulu, P. Sreenivasa Kumar, C. E. Veni Madhavan
1991 conf
FSTTCS
Subir Kumar Ghosh, Anil Maheshwari, Sudebkumar Prasant Pal, Sanjeev Saluja, C. E. Veni Madhavan
1990 conf
SCT
V. Vinay, H. Venkateswaran, C. E. Veni Madhavan
1990 ed.
FSTTCS
Kesav V. Nori, C. E. Veni Madhavan
1989 conf
FSTTCS
P. Sreenivasa Kumar, C. E. Veni Madhavan
1989 J jnl
Inf. Process. Lett.
Bhaskar DasGupta, C. E. Veni Madhavan
1989 ed.
FSTTCS
C. E. Veni Madhavan
1987 J jnl
Theor. Comput. Sci.
V. S. Lakshmanan, C. E. Veni Madhavan
1987 conf
FSTTCS
P. Shanti Sastry, N. Jayakumar, C. E. Veni Madhavan
1986 conf
FSTTCS
V. S. Lakshmanan, C. E. Veni Madhavan
1985 J jnl
Inf. Process. Manag.
Nageswara S. V. Rao, S. Sitharama Iyengar, C. E. Veni Madhavan
1985 conf
FSTTCS
V. S. Lakshmanan, C. E. Veni Madhavan
1984 conf
FSTTCS
C. E. Veni Madhavan
1984 conf
FSTTCS
V. S. Lakshmanan, N. Chandrasekharan, C. E. Veni Madhavan
1984 J jnl
Theor. Comput. Sci.
C. E. Veni Madhavan
1983 J jnl
IEEE Trans. Computers
C. E. Veni Madhavan, S. Krishna
1980 J jnl
ACM Trans. Database Syst.
V. Gopalakrishna, C. E. Veni Madhavan
redb/extractors/js_extractors/js_strings.py
← Index redb/extractors/js_extractors/js_strings.py python
import base64
import bisect
import inspect
import re
from datetime import datetime, timezone
from typing import Any

from redb.extractors.enum import Tag
from redb.extractors.js_extractor import JSExtractor
from redb.extractors.js_extractors.js_patterns import STRING_PATTERNS, line_offsets

# Local aliases for the compiled patterns this extractor uses. Defined and
# compiled exactly once in js_patterns.STRING_PATTERNS.
_HEX_STRING_RE = STRING_PATTERNS["hex_escape_seq"]
_UNICODE_STRING_RE = STRING_PATTERNS["unicode_escape_seq"]
_CHARCODE_RE = STRING_PATTERNS["charcode_call"]
_BASE64_STRING_RE = STRING_PATTERNS["base64_quoted"]
_CONCAT_STRING_RE = STRING_PATTERNS["concat_chain"]

# Tokeniser used inside _reconstruct_concat to pull each quoted part out of a
# matched concat chain. Compiled once at module load (was recompiled on every
# concat match before).
_CONCAT_TOKEN_RE = re.compile(r'["\']([^"\']*)["\']')


class JSStringsExtractor(JSExtractor):

    def __init__(
        self, filepath, log, exporters=None, index_prefix=None,
        known_benign=False, known_malicious=False, source=None, context=None,
    ):
        super().__init__(
            filepath, log, exporters, index_prefix,
            known_benign, known_malicious, source, context=context,
        )
        self.string_findings = None
        self.log.debug(inspect.currentframe().f_code.co_name)

    def tag(self):
        return Tag.JS_STRINGS.value

    def _decode_hex_string(self, hex_str):
        """Decode \\x41\\x42 style hex strings."""
        try:
            # Remove \\x prefix and decode
            clean = hex_str.replace('\\x', '')
            return bytes.fromhex(clean).decode('utf-8', errors='replace')
        except Exception:
            return None

    def _decode_unicode_string(self, uni_str):
        """Decode \\u0041\\u0042 style unicode strings."""
        try:
            return uni_str.encode('utf-8').decode('unicode_escape')
        except Exception:
            return None

    def _decode_charcode(self, charcode_str):
        """Decode String.fromCharCode(72, 101, 108, ...) sequences."""
        try:
            codes = [int(c.strip()) for c in charcode_str.split(',') if c.strip().isdigit()]
            return ''.join(chr(c) for c in codes if 0 <= c <= 0x10FFFF)
        except Exception:
            return None

    def _decode_base64(self, b64_str):
        """Attempt to decode base64 string."""
        try:
            decoded = base64.b64decode(b64_str)
            # Check if result is printable text
            text = decoded.decode('utf-8', errors='strict')
            # Only return if it looks like text (>80% printable)
            printable = sum(1 for c in text if c.isprintable() or c in '\n\r\t')
            if printable / len(text) > 0.8:
                return text
        except Exception:
            pass
        return None

    def _reconstruct_concat(self, concat_match):
        """Reconstruct concatenated string parts."""
        try:
            parts = _CONCAT_TOKEN_RE.findall(concat_match)
            return ''.join(parts)
        except Exception:
            return None

    def _find_line_number(self, match_start):
        """1-indexed line number for `match_start`, looked up in O(log L) via
        bisect over `self._line_offsets` (built once per extract() call).

        Replaces the historical `self.js_source[:match_start].count('\\n') + 1`
        which was O(N) per call and quadratic across all matches in a sample.
        """
        return bisect.bisect_right(self._line_offsets, match_start)

    def _scan_text(self, text):
        """Run every encoded-string pattern over `text` and return a list of
        finding dicts. Stateless apart from the per-call `_line_offsets` cache,
        which `_find_line_number` reads — callers must reset it before invoking
        this so line numbers reference the text being scanned, not the previous
        one.
        """
        findings = []

        # Hex-encoded strings
        for m in _HEX_STRING_RE.finditer(text):
            raw = m.group()
            decoded = self._decode_hex_string(raw)
            if decoded and len(decoded) >= 4:
                findings.append({
                    'string': decoded[:4000],
                    'string_raw': raw[:4000],
                    'string_encoding': 'hex',
                    'string_offset': self._find_line_number(m.start()),
                    'string_length': len(decoded),
                    'string_raw_length': len(raw),
                    'string_entropy': self._calculate_text_entropy(decoded),
                })

        # Unicode-encoded strings
        for m in _UNICODE_STRING_RE.finditer(text):
            raw = m.group()
            decoded = self._decode_unicode_string(raw)
            if decoded and len(decoded) >= 3:
                findings.append({
                    'string': decoded[:4000],
                    'string_raw': raw[:4000],
                    'string_encoding': 'unicode',
                    'string_offset': self._find_line_number(m.start()),
                    'string_length': len(decoded),
                    'string_raw_length': len(raw),
                    'string_entropy': self._calculate_text_entropy(decoded),
                })

        # String.fromCharCode sequences
        for m in _CHARCODE_RE.finditer(text):
            raw = m.group()
            decoded = self._decode_charcode(m.group(1))
            if decoded and len(decoded) >= 4:
                findings.append({
                    'string': decoded[:4000],
                    'string_raw': raw[:4000],
                    'string_encoding': 'charcode',
                    'string_offset': self._find_line_number(m.start()),
                    'string_length': len(decoded),
                    'string_raw_length': len(raw),
                    'string_entropy': self._calculate_text_entropy(decoded),
                })

        # Base64-encoded strings
        for m in _BASE64_STRING_RE.finditer(text):
            raw = m.group(0)
            b64_val = m.group(1)
            decoded = self._decode_base64(b64_val)
            if decoded and len(decoded) >= 10:
                findings.append({
                    'string': decoded[:4000],
                    'string_raw': raw[:4000],
                    'string_encoding': 'base64',
                    'string_offset': self._find_line_number(m.start()),
                    'string_length': len(decoded),
                    'string_raw_length': len(raw),
                    'string_entropy': self._calculate_text_entropy(decoded),
                })

        # Concatenated strings (reassembled)
        for m in _CONCAT_STRING_RE.finditer(text):
            raw = m.group()
            reconstructed = self._reconstruct_concat(raw)
            if reconstructed and len(reconstructed) >= 20:
                findings.append({
                    'string': reconstructed[:4000],
                    'string_raw': raw[:4000],
                    'string_encoding': 'concat',
                    'string_offset': self._find_line_number(m.start()),
                    'string_length': len(reconstructed),
                    'string_raw_length': len(raw),
                    'string_entropy': self._calculate_text_entropy(reconstructed),
                })

        return findings

    def extract(self):
        src = self.js_source
        if not src:
            return None

        # Pass 1: raw source. _line_offsets is keyed off whichever text is
        # currently being scanned so _find_line_number resolves to that text.
        self._line_offsets = line_offsets(src)
        findings = self._scan_text(src)

        # Pass 2: deobfuscated text, when the deobfuscator produced something
        # meaningfully different. Same patterns, but a different surface — for
        # samples where the encoded payload is hidden behind an outer wrapper
        # (e.g. array.join() + eval in Vjw0rm/WSH-RAT) only this pass yields
        # any rows at all.
        deobf_text, _ = self._context.deobfuscated
        if deobf_text and deobf_text != src:
            self._line_offsets = line_offsets(deobf_text)
            findings.extend(self._scan_text(deobf_text))

        if not findings:
            return None

        # Deduplicate by decoded string value (raw pass wins on collision: it
        # comes first in `findings`). A string that surfaces only in the
        # deobfuscated text still gets persisted, which is the whole point of
        # the second pass.
        seen_values = set()
        deduped = []
        for f in findings:
            val_key = f['string'][:100]
            if val_key not in seen_values:
                seen_values.add(val_key)
                deduped.append(f)

        self.string_findings = deduped[:500]  # Limit per file
        # Publish to the shared context so post-loop consumers (notably the IOC
        # plumbing in workers.py) can scrape the decoded strings without
        # holding a reference to this extractor instance.
        self._context.decoded_strings = self.string_findings
        return self.string_findings

    def prepare_export_data(self, exporter_type: str) -> Any:
        if exporter_type == "ClickHouseExporter":
            if not self.string_findings:
                return None

            data = []
            for f in self.string_findings:
                data.append([
                    self.sha256,
                    f['string'],
                    f['string_raw'],
                    f['string_encoding'],
                    f['string_offset'],
                    f['string_length'],
                    f['string_raw_length'],
                    f['string_entropy'],
                ])

            column_names = [
                "sha256",
                "string",
                "string_raw",
                "string_encoding",
                "string_offset",
                "string_length",
                "string_raw_length",
                "string_entropy",
            ]

            column_type_names = [
                "FixedString(64)",
                "String",
                "String",
                "LowCardinality(String)",
                "UInt64",
                "UInt32",
                "UInt32",
                "Float32",
            ]

            return (data, column_names, column_type_names)

    def get_clickhouse_table(self) -> str:
        return "code_binja_strings_raw"