Oliver Creighton

25 papers A 3B 3C 4Unranked 12
YearRankTypeTitle / Venue / Authors
2017 A conf
ICSA
Stefan Kugele, Philipp Obergfell, Manfred Broy, Oliver Creighton, Matthias Traub, Wolfgang Hopfensitz
2013 conf
Onward!
Han Xu, Oliver Creighton, Naoufel Boulila, Ruth Barbara Demmel
2013 A conf
RE
Oliver Creighton, Marcos R. S. Borges
2012 conf
SIGSOFT FSE
Han Xu, Oliver Creighton, Naoufel Boulila, Bernd Bruegge
2012 conf
RESS
Deepti Savio, P. C. Anitha, Arpith Patil, Oliver Creighton
2011 conf
MERE
Oliver Creighton, David Callele, Orlena Gotel
2011 conf
MERE
Steve Russell, Oliver Creighton
2011 conf
MERE
Harald Stangl, Oliver Creighton
2011 ed.
MERE
Oliver Creighton, David Callele, Orlena Gotel
2011 conf
Onward!
Naoufel Boulila, Oliver Creighton, Georgi A. Markov, Steve Russell, Ronald Blechner
2011 B conf
REFSQ
Georgi A. Markov, Anne Hoffmann, Oliver Creighton
2011 conf
MERE
Steve Russell, Oliver Creighton
2010 B conf
REFSQ
Benedikt Gleich, Oliver Creighton, Leonid Kof
2009 conf
OOPSLA Companion
Oliver Creighton, Ruth Barbara Demmel, Harald Stangl, Asa MacWilliams
2008 conf
MERE
Bernd Bruegge, Oliver Creighton, Maximilian Reiß, Harald Stangl
2008 B conf
SPLC
Anil Kumar Thurimella, Bernd Bruegge, Oliver Creighton
2008 conf
LMSA@ICSE
Oliver Creighton, Matthias Singer
2008 C conf
Software Engineering
Walid Maalej, Oliver Creighton, Ernst Pohn
2008 C conf
Software Engineering (Workshops)
Walid Maalej, Oliver Creighton, Ernst Pohn
2007 C conf
Software Engineering
Oliver Creighton, Bernd Bruegge
2006 ed.
MERE
Oliver Creighton, Bernd Bruegge
2006 A conf
RE
Oliver Creighton, Martin Ott, Bernd Brügge
2006
Oliver Creighton
2001 C conf
APSEC
Allen H. Dutoit, Oliver Creighton, Gudrun Klinker, Rafael Kobylinski, Christoph Vilsmeier, Bernd Brügge
2001 conf
ISAR
Gudrun Klinker, Oliver Creighton, Allen H. Dutoit, Rafael Kobylinski, Christoph Vilsmeier, Bernd Brügge
redb/extractors/js_extractors/js_content.py
← Index redb/extractors/js_extractors/js_content.py python
"""Persists raw + normalised text into the generic `code_text_content` table.

Reads the raw source and the deobfuscation result directly from the shared
JSContext so no extra compute happens here — both values are computed once
per sample (the source at JSContext construction, the deobfuscation lazily
on first access) and reused by any extractor that needs them.

`text_normalized` is left NULL when the deobfuscation pass produced no
output, so analysts can distinguish "we tried and got nothing" from
"normalisation succeeded".
"""

import inspect
from datetime import datetime, timezone
from typing import Any

from redb.extractors.enum import Tag
from redb.extractors.js_extractor import JSExtractor


class JSContentExtractor(JSExtractor):

    def __init__(
        self, filepath, log, exporters=None, index_prefix=None,
        known_benign=False, known_malicious=False, source=None, context=None,
    ):
        super().__init__(
            filepath, log, exporters, index_prefix,
            known_benign, known_malicious, source, context=context,
        )
        self.content_row = None
        self.log.debug(inspect.currentframe().f_code.co_name)

    def tag(self):
        return Tag.JS_CONTENT.value

    def extract(self):
        src = self.js_source
        if not src:
            return None

        deobfuscated, normalizer_used = self._context.deobfuscated

        self.content_row = {
            "content_type": self._context.content_type,
            "text_raw": src,
            "text_normalized": deobfuscated,  # may be None
            "normalizer_used": normalizer_used,  # may be None
        }
        return self.content_row

    def prepare_export_data(self, exporter_type: str) -> Any:
        if exporter_type != "ClickHouseExporter":
            return None
        if not self.content_row:
            return None

        r = self.content_row
        data = [[
            self.sha256,
            r["content_type"],
            r["text_raw"],
            r["text_normalized"],
            r["normalizer_used"],
            datetime.now(timezone.utc),
        ]]

        column_names = [
            "sha256",
            "content_type",
            "text_raw",
            "text_normalized",
            "normalizer_used",
            "analysis_date",
        ]

        column_type_names = [
            "FixedString(64)",
            "LowCardinality(String)",
            "String",
            "Nullable(String)",
            "Nullable(String)",
            "DateTime64(3, 'UTC')",
        ]

        return (data, column_names, column_type_names)

    def get_clickhouse_table(self) -> str:
        return "code_text_content"