James M. Willenbring

21 papers C 2Journal 15Unranked 4
YearRankTypeTitle / Venue / Authors
2025 J jnl
CoRR
Lois Curfman McInnes, Dorian Arnold, Prasanna Balaprakash, Mike Bernhardt, Beth Cerny, Anshu Dubey, Roscoe Giles, Denice Ward Hood, Mary Ann E. Leung, Vanessa Lopez-Marrero, Paul Messina, Olivia B. Newton, Chris Oehmen, Stefan M. Wild, James M. Willenbring, Lou Woodley, Tony Baylis, David E. Bernholdt, Chris Camano, Johanna Cohoon, Charles Ferenbaugh, Stephen M. Fiore, Sandra Gesing, Diego Gómez-Zará, James Howison, Tanzima Z. Islam, David Kepczynski, Charles Lively, Harshitha Menon, Bronson Messer, Marieme Ngom, Umesh Paliath, Michael E. Papka, Irene Qualters, Elaine M. Raybourn, Katherine Riley, Paulina Rodriguez, Damian W. I. Rouson, Michelle Schwalbe, Sudip K. Seal, Ozge Surer, Valerie Taylor, Lingfei Wu
2025 J jnl
CoRR
Matthias Mayr, Alexander Heinlein, Christian A. Glusa, Siva Rajamanickam, Maarten Arnst, Roscoe A. Bartlett, Luc Berger-Vergiat, Erik G. Boman, Karen D. Devine, Graham Harper, Michael A. Heroux, Mark Hoemmen, Jonathan J. Hu, Brian Michael Kelley, Kyungjoo Kim, Drew P. Kouri, Paul Kuberry, Kim Liegeois, Curtis C. Ober, Roger P. Pawlowski, Carl Pearson, Mauro Perego, Eric T. Phipps, Denis Ridzal, Nathan V. Roberts, Christopher M. Siefert, Heidi Thornquist, Romin Tomasetti, Christian R. Trott, Raymond S. Tuminaro, James M. Willenbring, Michael M. Wolf, Ichitaro Yamazaki
2024 J jnl
Comput. Sci. Eng.
Lois Curfman McInnes, Michael A. Heroux, David E. Bernholdt, Anshu Dubey, Elsa Gonsiorowski, Rinku Gupta, Osni Marques, J. David Moulton, Hai Ah Nam, Boyana Norris, Elaine M. Raybourn, James M. Willenbring, Ann S. Almgren, Roscoe A. Bartlett, Kita Cranfill, Stephen Fickas, Don Frederick, William F. Godoy, Patricia A. Grubel, Rebecca Hartman-Baker, Axel Huebl, Rose Lynch, Addi Malviya-Thakur, Reed Milewicz, Mark C. Miller, Miranda Mundt, Erik Palmer, Suzanne Parete-Koon, Megan Phinney, Katherine Riley, David M. Rogers, Benjamin H. Sims, Deborah Stevens, Gregory R. Watson
2024 J jnl
Comput. Sci. Eng.
James M. Willenbring, Sameer Suresh Shende, Todd Gamblin
2024 J jnl
Future Gener. Comput. Syst.
James M. Willenbring, Gursimran Singh Walia
2024 J jnl
CoRR
Michael A. Heroux, Sameer Shende, Lois Curfman McInnes, Todd Gamblin, James M. Willenbring
2023 J jnl
CoRR
Lois Curfman McInnes, Michael A. Heroux, David E. Bernholdt, Anshu Dubey, Elsa Gonsiorowski, Rinku Gupta, Osni Marques, J. David Moulton, Hai Ah Nam, Boyana Norris, Elaine M. Raybourn, James M. Willenbring, Ann S. Almgren, Ross Bartlett, Kita Cranfill, Stephen Fickas, Don Frederick, William F. Godoy, Patricia Grubel, Rebecca Hartman-Baker, Axel Huebl, Rose Lynch, Addi Malviya-Thakur, Reed Milewicz, Mark C. Miller, Miranda Mundt, Erik Palmer, Suzanne Parete-Koon, Megan Phinney, Katherine Riley, David M. Rogers, Benjamin H. Sims, Deborah Stevens, Gregory R. Watson
2022 J jnl
Comput. Sci. Eng.
Cody J. Balos, Piotr Luszczek, Sarah Osborn, James M. Willenbring, Ulrike Meier Yang
2022 C conf
SEKE
James M. Willenbring, Gursimran Singh Walia
2022 conf
ISSRE Workshops
James M. Willenbring, Gursimran Singh Walia
2020 J jnl
CoRR
Reed Milewicz, James M. Willenbring, Dena Vigil
2019 conf
HUST/SE-HER/WIHPC@SC
Michael A. Heroux, Elsa Gonsiorowski, Rinku Gupta, Reed Milewicz, J. David Moulton, Gregory R. Watson, James M. Willenbring, Richard J. Zamora, Elaine M. Raybourn
2017 J jnl
CoRR
Roscoe A. Bartlett, Irina Demeshko, Todd Gamblin, Glenn Hammond, Michael A. Heroux, Jeffrey Johnson, Alicia M. Klinvex, Xiaoye S. Li, Lois Curfman McInnes, J. David Moulton, Daniel Osei-Kuffuor, Jason Sarich, Barry Smith, James M. Willenbring, Ulrike Meier Yang
2017 J jnl
Supercomput. Front. Innov.
Roscoe A. Bartlett, Irina Demeshko, Todd Gamblin, Glenn Hammond, Michael A. Heroux, Jeffrey Johnson, Alicia M. Klinvex, Xiaoye S. Li, Lois Curfman McInnes, J. David Moulton, Daniel Osei-Kuffuor, Jason Sarich, Barry Smith, James M. Willenbring, Ulrike Meier Yang
2015 J jnl
ACM Trans. Math. Softw.
James M. Willenbring
2013 J jnl
CoRR
Chetan Jhurani, Travis M. Austin, Michael A. Heroux, James M. Willenbring
2012 J jnl
Sci. Program.
Michael A. Heroux, James M. Willenbring
2012 conf
eScience
Roscoe A. Bartlett, Michael A. Heroux, James M. Willenbring
2009 conf
SE-CSE@ICSE
Michael A. Heroux, James M. Willenbring
2007 C conf
PDP
Michael A. Heroux, James M. Willenbring, Michael N. Phenow
2005 J jnl
ACM Trans. Math. Softw.
Michael A. Heroux, Roscoe A. Bartlett, Vicki E. Howle, Robert J. Hoekstra, Jonathan J. Hu, Tamara G. Kolda, Richard B. Lehoucq, Kevin R. Long, Roger P. Pawlowski, Eric T. Phipps, Andrew G. Salinger, Heidi Thornquist, Ray S. Tuminaro, James M. Willenbring, Alan B. Williams, Kendall S. Stanley
redb/extractors/js_extractor.py
← Index redb/extractors/js_extractor.py python
import logging
import re
from abc import ABCMeta, abstractmethod

from redb.extractors.extractor import Extractor
from redb.extractors.js_extractors.js_context import JSContext, _text_entropy

logger = logging.getLogger(__name__)

# ESM is recognised by line-anchored `import ... from "..."` / bare side-effect
# `import "..."` / top-level `export ...`. Anchored at line start to avoid
# matching the substring inside string literals or comments.
_ESM_PATTERN = re.compile(
    r'(?m)^\s*(?:'
    r'import\s+[^;\n]*?\bfrom\s+[\'"]'
    r'|import\s+[\'"][^\'"]+[\'"]'
    r'|export\s+(?:default\b|\{|\*|const\b|let\b|var\b|function\b|class\b|async\b)'
    r')'
)


@abstractmethod
class JSExtractor(Extractor, metaclass=ABCMeta):
    """Base class for JavaScript file extractors.

    Every JSExtractor reads its raw materials (bytes / decoded source / line
    list / scan_source results / pyjsparser AST / text entropy) from a shared
    `JSContext`. When workers.py drives the JS pipeline it builds one context
    per sample and threads it into every extractor via `context=`. When tests
    or other callers instantiate an extractor directly, the constructor builds
    a fresh context from `(filepath, source=...)`.

    All historical instance attributes (`self.binary`, `self.js_source`,
    `self.lines`) and helpers (`self._decode_source`, `self._parse_ast`,
    `self._calculate_text_entropy`) are preserved as thin delegators so
    existing extractor code keeps working unchanged.
    """

    def __init__(
        self,
        filepath,
        log,
        exporters=None,
        index_prefix=None,
        known_benign=False,
        known_malicious=False,
        source=None,
        context=None,
    ):
        if context is None:
            context = JSContext.from_path(filepath, log=log, source=source)
        elif source is not None and context.source != source:
            log.warning(
                "JSExtractor received both `source=` and `context=` with "
                "differing source; ignoring source kwarg"
            )
        self._context = context

        super().__init__(
            filepath,
            log,
            exporters,
            index_prefix,
            known_benign=known_benign,
            known_malicious=known_malicious,
        )

    @property
    def binary(self):
        return self._context.raw_bytes

    @property
    def js_source(self):
        return self._context.source

    @property
    def lines(self):
        return self._context.lines

    def _decode_source(self):
        """Back-compat shim — the context already decoded once at construction.

        Kept so any external caller using the historical method name keeps
        working without touching the underlying bytes again.
        """
        return self._context.source

    def _parse_ast(self):
        """Return the shared pyjsparser AST (or None if unavailable)."""
        return self._context.ast

    def _calculate_text_entropy(self, text):
        """Shannon text entropy for `text`.

        When `text` is the context's own source we read the cached value;
        otherwise we compute fresh. JSStringsExtractor calls this on arbitrary
        decoded substrings, so the fresh-compute path must remain available.
        """
        if text is self._context.source:
            return self._context.text_entropy
        return _text_entropy(text)

    def _detect_environment(self):
        """Detect the target JS runtime environment."""
        src = self.js_source
        if not src:
            return "unknown"

        # WScript/WSH indicators
        wscript_patterns = [
            'WScript.', 'WSH.', 'ActiveXObject', 'Scripting.FileSystemObject',
            'WScript.Shell', 'ADODB.Stream',
        ]
        for p in wscript_patterns:
            if p in src:
                return "wscript"

        # Browser-extension APIs — checked before generic browser/worker because
        # `chrome.*` and `browser.runtime` are distinctive of MV2/MV3 extensions
        extension_patterns = [
            'chrome.runtime', 'chrome.tabs', 'chrome.storage',
            'chrome.webRequest', 'browser.runtime', 'browser.tabs',
        ]
        for p in extension_patterns:
            if p in src:
                return "browser_extension"

        # Service / Web Workers — worker-only APIs that don't appear in regular
        # browser pages (a generic browser script would use `window.` or
        # `document.`, never `self.importScripts` or `caches.match`)
        worker_patterns = [
            "self.addEventListener('fetch'", 'self.addEventListener("fetch"',
            'self.importScripts', 'self.skipWaiting',
            'caches.match', 'caches.open',
        ]
        for p in worker_patterns:
            if p in src:
                return "service_worker"

        # Deno runtime
        if 'Deno.' in src:
            return "deno"

        # Node.js indicators
        node_patterns = [
            'require(', 'module.exports', 'process.env', '__dirname',
            '__filename', 'Buffer.', 'child_process',
        ]
        for p in node_patterns:
            if p in src:
                return "node"

        # Browser indicators
        browser_patterns = [
            'document.', 'window.', 'navigator.', 'localStorage',
            'sessionStorage', 'XMLHttpRequest', 'addEventListener',
        ]
        for p in browser_patterns:
            if p in src:
                return "browser"

        return "unknown"

    def _detect_script_type(self):
        """Detect the script type/format."""
        src = self.js_source
        if not src:
            return "unknown"

        stripped = src.lstrip()

        # JScript.Encode payload — must be checked first since the encoded
        # body can't be classified any other way
        if stripped.startswith('#@~^'):
            return "jse"

        # WSF / HTA live in the first few KB of an HTML-ish wrapper
        head_lower = stripped[:4096].lower()

        # Windows Script File — XML wrapper around one or more <script> blocks
        if ('<job' in head_lower or '<package' in head_lower) and '<script' in head_lower:
            return "wsf"

        # HTML Application — distinct from generic embedded_html because HTAs
        # run under mshta.exe with full WSH/ActiveX access
        if '<hta:application' in head_lower or 'application/hta' in head_lower:
            return "hta"

        if stripped.startswith('<!') or stripped.startswith('<html') or '<script' in stripped[:2000]:
            return "embedded_html"

        if 'WScript.' in src or 'WSH.' in src:
            return "wscript"

        if _ESM_PATTERN.search(src):
            return "esm"

        if 'require(' in src or 'module.exports' in src:
            return "node_module"

        return "standalone"