James W. Mickens

22 papers A* 6A 3B 1Misc 4Journal 4Unranked 3
YearRankTypeTitle / Venue / Authors
2021 A* conf
USENIX Security Symposium
Xueyuan Han, Xiao Yu, Thomas F. J.-M. Pasquier, Ding Li, Junghwan Rhee, James W. Mickens, Margo I. Seltzer, Haifeng Chen
2014 Misc conf
NSDI
James W. Mickens, Edmund B. Nightingale, Jeremy Elson, Darren Gehring, Bin Fan, Asim Kadav, Vijay Chidambaram, Osama Khan, Krishna Nareddy
2013 A conf
FAST
Jacob R. Lorch, Bryan Parno, James W. Mickens, Mariana Raykova, Joshua Schiffman
2012 J jnl
IACR Cryptol. ePrint Arch.
Jacob R. Lorch, James W. Mickens, Bryan Parno, Mariana Raykova, Joshua Schiffman
2011 A* conf
SOSP
James W. Mickens, Mohan Dhawan
2011 A* conf
IEEE Symposium on Security and Privacy
Bryan Parno, Jacob R. Lorch, John R. Douceur, James W. Mickens, Jonathan M. McCune
2010 A* conf
INFOCOM
John R. Douceur, James W. Mickens, Thomas Moscibroda, Debmalya Panigrahi
2010 Misc conf
NSDI
James W. Mickens, Jeremy Elson, Jon Howell, Jay R. Lorch
2010 Misc conf
NSDI
James W. Mickens, Jeremy Elson, Jon Howell
2009 A* conf
PODC
John R. Douceur, James W. Mickens, Thomas Moscibroda, Debmalya Panigrahi
2009 J jnl
ACM SIGOPS Oper. Syst. Rev.
James W. Mickens, Dilma Da Silva
2009 conf
USENIX ATC
James W. Mickens, John R. Douceur, William J. Bolosky, Brian D. Noble
2009 A conf
CoNEXT
John R. Douceur, James W. Mickens, Thomas Moscibroda, Debmalya Panigrahi
2008
James W. Mickens
2007 B conf
WiMob
James W. Mickens, Brian D. Noble
2007 A conf
DSN
James W. Mickens, Brian D. Noble
2006 Misc conf
NSDI
James W. Mickens, Brian D. Noble
2006 J jnl
SIGMETRICS Perform. Evaluation Rev.
James W. Mickens, Brian D. Noble
2005 conf
Workshop on Wireless Security
James W. Mickens, Brian D. Noble
2005 A* conf
SIGMETRICS
James W. Mickens, Brian D. Noble
2004 conf
Secure Data Management
Magesh Jayapandian, Brian D. Noble, James W. Mickens, H. V. Jagadish
2003 J jnl
IEEE Pervasive Comput.
James W. Mickens, Marco Gruteser, Justin Mazzola Paluska, Mary Baker
redb/extractors/js_extractors/js_content.py
← Index redb/extractors/js_extractors/js_content.py python
"""Persists raw + normalised text into the generic `code_text_content` table.

Reads the raw source and the deobfuscation result directly from the shared
JSContext so no extra compute happens here — both values are computed once
per sample (the source at JSContext construction, the deobfuscation lazily
on first access) and reused by any extractor that needs them.

`text_normalized` is left NULL when the deobfuscation pass produced no
output, so analysts can distinguish "we tried and got nothing" from
"normalisation succeeded".
"""

import inspect
from datetime import datetime, timezone
from typing import Any

from redb.extractors.enum import Tag
from redb.extractors.js_extractor import JSExtractor


class JSContentExtractor(JSExtractor):

    def __init__(
        self, filepath, log, exporters=None, index_prefix=None,
        known_benign=False, known_malicious=False, source=None, context=None,
    ):
        super().__init__(
            filepath, log, exporters, index_prefix,
            known_benign, known_malicious, source, context=context,
        )
        self.content_row = None
        self.log.debug(inspect.currentframe().f_code.co_name)

    def tag(self):
        return Tag.JS_CONTENT.value

    def extract(self):
        src = self.js_source
        if not src:
            return None

        deobfuscated, normalizer_used = self._context.deobfuscated

        self.content_row = {
            "content_type": self._context.content_type,
            "text_raw": src,
            "text_normalized": deobfuscated,  # may be None
            "normalizer_used": normalizer_used,  # may be None
        }
        return self.content_row

    def prepare_export_data(self, exporter_type: str) -> Any:
        if exporter_type != "ClickHouseExporter":
            return None
        if not self.content_row:
            return None

        r = self.content_row
        data = [[
            self.sha256,
            r["content_type"],
            r["text_raw"],
            r["text_normalized"],
            r["normalizer_used"],
            datetime.now(timezone.utc),
        ]]

        column_names = [
            "sha256",
            "content_type",
            "text_raw",
            "text_normalized",
            "normalizer_used",
            "analysis_date",
        ]

        column_type_names = [
            "FixedString(64)",
            "LowCardinality(String)",
            "String",
            "Nullable(String)",
            "Nullable(String)",
            "DateTime64(3, 'UTC')",
        ]

        return (data, column_names, column_type_names)

    def get_clickhouse_table(self) -> str:
        return "code_text_content"