Oscar Karnalim

66 papers A 2B 10C 6Journal 26Unranked 22
YearRankTypeTitle / Venue / Authors
2026 J jnl
Discov. Comput.
Oscar Karnalim, Hapnes Toba, Maresha Caroline Wijanto, Yehezkiel David Setiawan, Nico Surantha, Michael Liut
2026 J jnl
Softw. Impacts
Oscar Karnalim, Yehezkiel David Setiawan
2025 B conf
ICALT
Oscar Karnalim, Adelia, Diana Trivena Yulianti, Doro Edi, Judea J. Jarden
2025 B conf
COMPSAC
Suqing Liu, Lisa Zhang, Oscar Karnalim, Michael Liut
2025 conf
CompEd (1)
Caroline Pechenik, Angela M. Zavaleta Bernuy, Selina Marianna Shah, Shirley de Wit, Emmanuel Awuni Kolog, Oscar Karnalim, Mohammed Farghally, Carlos Aníbal Suárez, Jack Parkinson, Leo Porter, Rodrigo Duran, Paul Vrbik, Brian Harrington, Lisa Zhang, Michael Liut, Andrew Petersen
2025 B conf
ICALT
Oscar Karnalim, Yehezkiel David Setiawan, Rossevine Artha Nathasya, Femmy Friscilla Susilo
2024 conf
ICEEL
Hapnes Toba, Laurentius Gusti Ontoseno Panata Yudha, Oscar Karnalim, Hendra Bunyamin, Terutoshi Tada
2024 conf
SIGCSE Virtual (2)
Erik Barendsen, Violetta Lonati, Keith Quille, Rukiye Altin, Monica Divitini, Sara Hooshangi, Oscar Karnalim, Natalie Kiesler, Madison Melton, Calkin Suero Montero, Anna Morpurgo
2024 J jnl
Educ. Inf. Technol.
Oscar Karnalim, Hapnes Toba, Meliana Christianti Johan
2024 B conf
COMPSAC
Michael Sheinman Orenstrakh, Oscar Karnalim, Carlos Aníbal Suárez, Michael Liut
2024 B conf
ICALT
Oscar Karnalim, Mewati Ayub, Krismanto Kusbiantoro
2024 J jnl
Comput. Sci. Educ.
Oscar Karnalim, Simon, William J. Chivers
2024 J jnl
Informatics Educ.
Oscar Karnalim
2024 conf
SIGCSE Virtual (Working Group Reports)
Erik Barendsen, Violetta Lonati, Keith Quille, Rukiye Altin, Monica Divitini, Sara Hooshangi, Oscar Karnalim, Natalie Kiesler, Madison Melton, Calkin Suero Montero, Anna Morpurgo
2024 conf
MIPRO
Mewati Ayub, Oscar Karnalim
2023 J jnl
CoRR
Michael Sheinman Orenstrakh, Oscar Karnalim, Carlos Aníbal Suárez, Michael Liut
2023 J jnl
IEEE Trans. Learn. Technol.
Oscar Karnalim, Simon, William J. Chivers
2023 B conf
ICALT
Oscar Karnalim
2023 J jnl
CoRR
Hapnes Toba, Oscar Karnalim, Meliana Christianti Johan, Terutoshi Tada, Yenni Merlin Djajalaksana, Tristan Vivaldy
2023 conf
TALE
Oscar Karnalim, Erico Darmawan Handoyo, Hapnes Toba, Yehezkiel David Setiawan, Meliana Christianti Johan, Josephine Alvina Luwia
2023 J jnl
CoRR
Oscar Karnalim, Hapnes Toba, Meliana Christianti Johan, Erico Darmawan Handoyo, Yehezkiel David Setiawan, Josephine Alvina Luwia
2022 conf
WCCE
Oscar Karnalim, Simon, William J. Chivers, Billy Susanto Panca
2022 J jnl
ACM Trans. Comput. Educ.
Oscar Karnalim, Simon, William J. Chivers, Billy Susanto Panca
2022 J jnl
Comput. Appl. Eng. Educ.
Oscar Karnalim, Simon, William J. Chivers
2022 conf
ITiCSE (2)
Jeremiah J. Blanchard, John R. Hott, Vincent Berry, Rebecca Carroll, Bob Edmison, Richard Glassey, Oscar Karnalim, Brian Plancher, Seán Russell
2022 conf
WCCE
Muftah Afrizal Pangestu, Simon, Oscar Karnalim
2022 conf
ICL (1)
Oscar Karnalim, Simon, William J. Chivers
2022 B conf
ICALT
Oscar Karnalim, Afifah Muharikah, Sunarto Natsir
2022 conf
ITiCSE-WGR
Jeremiah J. Blanchard, John R. Hott, Vincent Berry, Rebecca Carroll, Bob Edmison, Richard Glassey, Oscar Karnalim, Brian Plancher, Seán Russell
2022 C conf
EDUCON
Oscar Karnalim, Simon, William J. Chivers
2022 conf
ICL (2)
Mewati Ayub, Oscar Karnalim, Maresha Caroline Wijanto, Robby Tan, Risal, Rossevine Artha Nathasya
2021 A conf
ICER
Oscar Karnalim
2021 A conf
SIGCSE
Oscar Karnalim, Simon
2021 J jnl
IEEE Access
Oscar Karnalim, Simon
2021 conf
TALE
Oscar Karnalim, Gisela Kurniawati, Sendy Ferdian Suiadi
2021 conf
TALE
Oscar Karnalim, Irwan Alnarus Kautsar, Bayu Rima Aditya, Yogi Udjaja, Matahari Bhakti Nendya, I Nyoman Darma Kotama
2021 C conf
FIE
Oscar Karnalim, Simon
2021 B conf
ICALT
Oscar Karnalim, Simon
2021 C conf
EDUCON
Maresha Caroline Wijanto, Oscar Karnalim, Mewati Ayub, Hapnes Toba, Robby Tan
2021 B conf
ICALT
Oscar Karnalim, Maresha Caroline Wijanto
2021 C conf
EDUCON
Oscar Karnalim, Simon, Mewati Ayub, Gisela Kurniawati, Rossevine Artha Nathasya, Maresha Caroline Wijanto
2020 J jnl
Int. J. Online Biomed. Eng.
Ricardo Franclinton, Oscar Karnalim, Mewati Ayub
2020 J jnl
Int. J. Online Biomed. Eng.
Billy Susanto Panca, Yansen Paulus, Oscar Karnalim
2020 conf
ITiCSE-WGR
Simon, Oscar Karnalim, Judy Sheard, Ilir Dema, Amey Karkare, Juho Leinonen, Michael Liut, Renée McCauley
2020 J jnl
Int. J. Eng. Pedagog.
Oscar Karnalim, Gisela Kurniawati, Sendy Ferdian Sujadi, Rossevine Artha Nathasya
2020 conf
Koli Calling
Oscar Karnalim, Simon
2020 conf
Koli Calling
Oscar Karnalim, Simon, William J. Chivers
2020 J jnl
Int. J. Comput.
Oscar Karnalim, Gisela Kurniawati
2020 B conf
ITiCSE
Simon, Oscar Karnalim, Judy Sheard, Ilir Dema, Amey Karkare, Juho Leinonen, Michael Liut, Renée McCauley
2020 C conf
ACE
Oscar Karnalim, Simon
2020 J jnl
Comput. Sci.
Oscar Karnalim
2019 J jnl
Comput.
Ariel Elbert Budiman, Oscar Karnalim
2019 J jnl
Comput. Appl. Eng. Educ.
Lisan Sulistiani, Oscar Karnalim
2019 J jnl
J. King Saud Univ. Comput. Inf. Sci.
Oscar Karnalim
2019 conf
TALE
Oscar Karnalim, Simon, William J. Chivers
2019 J jnl
Informatics Educ.
Oscar Karnalim, Setia Budi, Hapnes Toba, Mike Joy
2018 J jnl
Int. J. Online Eng.
Oscar Karnalim, Mewati Ayub
2018 J jnl
CoRR
Oscar Karnalim, Lisan Sulistiani
2018 conf
IPTA
Oscar Karnalim, Setia Budi, Sulaeman Santoso, Erico D. Handoyo, Hapnes Toba, Huyen Nguyen, Vishv Malhotra
2018 C conf
ISM
Setia Budi, Oscar Karnalim, Erico D. Handoyo, Sulaeman Santoso, Hapnes Toba, Huyen Nguyen, Vishv Malhotra
2018 J jnl
CoRR
Oscar Karnalim, Setia Budi
2018 conf
iCAST
Oscar Karnalim, Lisan Sulistiani
2018 J jnl
CoRR
Oscar Karnalim, Lisan Sulistiani
2017 J jnl
CoRR
Oscar Karnalim
2017 conf
ICMHI
Aulia Zahrina Qashri, Oscar Karnalim, Hapnes Toba
2016 conf
ACOMP
Oscar Karnalim
redb/extractors/js_extractors/js_context.py
← Index redb/extractors/js_extractors/js_context.py python
"""Per-sample shared state for the JavaScript extractor pipeline.

A `JSContext` is built exactly once per JS sample (in `workers.py`) and threaded
into every extractor that runs against that sample. It owns the disk read, the
decoded source text, the line-split cache, the Shannon text-entropy figure, the
shared `scan_source()` results, and the pyjsparser AST. Each of those is
computed lazily through `cached_property` so an extractor that doesn't need a
particular artefact does not pay for it.

Without this object, every JS extractor instance redoes the same disk read,
decode, scan, and (for any consumer) AST parse. With it, every extractor
shares one set of results.

`JSExtractor.__init__` accepts the context via a `context=` kwarg; if absent
(e.g. unit tests instantiating an extractor directly with `source=...`) it
builds a fresh context from the constructor arguments. Either path produces a
fully-populated context, so extractor code can always rely on
`self._context.scan` / `self._context.ast` / etc.
"""

from __future__ import annotations

import math
from collections import Counter
from dataclasses import dataclass
from functools import cached_property
from typing import Any, Dict, List, Optional

import chardet

from redb.extractors.js_extractors.js_patterns import scan_source


def decode_source(raw_bytes: bytes) -> str:
    """Decode raw JS bytes to text, honouring BOMs and falling back to chardet.

    Mirrors the historical `JSExtractor._decode_source` logic so existing tests
    continue to round-trip identically.
    """
    if not raw_bytes:
        return ""

    if raw_bytes[:3] == b"\xef\xbb\xbf":
        return raw_bytes[3:].decode("utf-8", errors="replace")
    if raw_bytes[:2] in (b"\xff\xfe", b"\xfe\xff"):
        return raw_bytes.decode("utf-16", errors="replace")

    try:
        return raw_bytes.decode("utf-8")
    except UnicodeDecodeError:
        pass

    try:
        detected = chardet.detect(raw_bytes)
        if detected and detected.get("encoding"):
            return raw_bytes.decode(detected["encoding"], errors="replace")
    except Exception:
        pass

    return raw_bytes.decode("latin-1", errors="replace")


def _text_entropy(text: str) -> float:
    """Shannon entropy of the character distribution of `text`, rounded to 4dp."""
    if not text:
        return 0.0
    counter = Counter(text)
    length = len(text)
    entropy = 0.0
    for count in counter.values():
        p = count / length
        if p > 0:
            entropy -= p * math.log2(p)
    return round(entropy, 4)


@dataclass
class JSContext:
    """Shared raw materials for one JS sample, consumed by every JS extractor.

    Cheap attributes (raw_bytes, source) are populated eagerly by the factory.
    Expensive ones (scan, ast) are cached_property — computed on first access
    and reused across every extractor that holds the same context.

    `content_type` is the magika label (e.g. `"javascript"`) carried alongside
    the source so the new code_text_content writer (and any future generic
    text-content writer) can record it without re-running magika. Defaults to
    `"javascript"` because by construction this context type is JS-specific;
    workers.py supplies the actual magika value when it builds the context.
    """

    filepath: str
    raw_bytes: bytes
    source: str
    log: Any = None
    content_type: str = "javascript"
    # Populated by JSStringsExtractor.extract() (the decoded/reconstructed
    # strings — hex/unicode/charcode/base64/concat unpacked into plaintext).
    # Read post-loop by the IOC plumbing in workers.py so any IOCs hidden
    # behind those encodings get scraped from the decoded form. Stays None
    # if JSStringsExtractor didn't run for this sample.
    decoded_strings: Optional[list] = None

    @cached_property
    def lines(self) -> List[str]:
        return self.source.splitlines() if self.source else []

    @cached_property
    def text_entropy(self) -> float:
        return _text_entropy(self.source)

    @cached_property
    def scan(self) -> Dict[str, Dict[str, object]]:
        """Result of running scan_source() exactly once over self.source."""
        return scan_source(self.source) if self.source else {}

    @cached_property
    def ast(self) -> Optional[Any]:
        """Lazy pyjsparser AST. Returns None if the parser is missing or fails.

        Extractors should treat None AST as "fall back to regex" — every
        AST-consuming extractor already handles that path.
        """
        if not self.source:
            return None
        try:
            import pyjsparser
            return pyjsparser.parse(self.source)
        except ImportError:
            if self.log is not None:
                self.log.debug("pyjsparser not installed, AST analysis skipped")
        except Exception as e:
            if self.log is not None:
                self.log.warning(f"AST parsing failed for {self.filepath}: {e}")
        return None

    @cached_property
    def deobfuscated(self) -> "tuple[Optional[str], Optional[str]]":
        """Run the configured JS deobfuscator (with jsbeautifier fallback) once
        per sample and cache the result. Returns `(text, normalizer_used)` or
        `(None, None)` if neither path produced output.

        Computed lazily on first access — samples whose pipeline never reads
        this don't pay the subprocess cost.
        """
        from redb.extractors.js_extractors.js_deobfuscator import deobfuscate
        return deobfuscate(self.source, self.log)

    @cached_property
    def scan_deobfuscated(self) -> Dict[str, Dict[str, object]]:
        """Result of running scan_source() exactly once over the deobfuscated
        text, keyed by PATTERNS only (FEATURE_PATTERNS are not consulted by
        the dual-pass consumers). Empty dict when there is no deobfuscated
        text or it equals the raw source.

        Two extractors consume the post-deobf API surface:
        `JSSuspiciousAPIsExtractor` (for revealed_by_deobf rows) and
        `JSDeobfuscationExtractor` (for the new_apis_found diff). Caching here
        means we scan the deobfuscated text once instead of twice per sample.
        """
        from redb.extractors.js_extractors.js_patterns import PATTERNS
        deobf_text, _ = self.deobfuscated
        if not deobf_text or deobf_text == self.source:
            return {}
        return scan_source(deobf_text, patterns=(PATTERNS,))

    @cached_property
    def xray(self):
        """Run @nodesecure/js-x-ray once per sample and cache the result.

        Returns an `XRayResult` (always — the function collapses every failure
        path to an empty result so callers don't have to special-case missing
        Node, missing package, timeouts, or parse errors). The
        `JSFeaturesExtractor` reads it for the obfuscator family name and for
        corroborating warning kinds; the heuristic falls back cleanly when
        `xray.obfuscator is None`.
        """
        from redb.extractors.js_extractors.js_xray import run
        return run(self.source, self.log)

    @classmethod
    def from_path(
        cls,
        filepath: str,
        log: Any = None,
        source: Optional[str] = None,
        raw_bytes: Optional[bytes] = None,
        content_type: str = "javascript",
    ) -> "JSContext":
        """Build a context from disk. `raw_bytes` and `source` are optional
        overrides — useful when the caller has already read or decoded the file.
        `content_type` is the magika label workers.py dispatched on; it lands
        on the context for the code_text_content writer to record.
        """
        if raw_bytes is None:
            with open(filepath, "rb") as f:
                raw_bytes = f.read()
        if source is None:
            source = decode_source(raw_bytes)
        return cls(
            filepath=filepath,
            raw_bytes=raw_bytes,
            source=source,
            log=log,
            content_type=content_type,
        )