Xiaodong Cheng

68 papers C 3Journal 38Unranked 27
YearRankTypeTitle / Venue / Authors
2026 J jnl
IEEE Trans. Syst. Man Cybern. Syst.
Ning Zhou, Canyang Zhao, Xiaodong Cheng, Yuanqing Xia, Tiejun Li, Baohao Wang
2025 J jnl
IEEE Internet Things J.
Hualong Chen, Yuanqiao Wen, Xiaodong Cheng, Changshi Xiao
2025 J jnl
IEEE Trans. Syst. Man Cybern. Syst.
Ning Zhou, Canyang Zhao, Xiaodong Cheng, Yuanqing Xia
2025 J jnl
CoRR
Giovanni Pugliese Carratelli, Xiaodong Cheng, Kris V. Parag, Ioannis Lestas
2025 J jnl
IEEE Trans. Autom. Control.
Xiaodong Cheng, Shengling Shi, Ioannis Lestas, Paul M. J. Van den Hof
2025 J jnl
IEEE Trans. Circuits Syst. II Express Briefs
Ning Zhou, Jialing Yan, Yuanqing Xia, Tiejun Li, Xiaodong Cheng
2024 conf
CCTA
Sjoerd Boersma, Xiaodong Cheng
2024 J jnl
Comput. Electr. Eng.
Xiaodong Cheng, Zhongyi Sui, Yuanqiao Wen, Dong Han
2024 J jnl
Comput. Electron. Agric.
Jan Lorenz Svensen, Xiaodong Cheng, Sjoerd Boersma, Congcong Sun
2024 conf
ECC
Mahdieh Sadat Sadabadi, Xiaodong Cheng
2024 conf
CDC
Yangming Dou, Xiaodong Cheng, Jacquelien M. A. Scherpen
2024 conf
CDC
Giovanni Pugliese Carratelli, Xiaodong Cheng, Kris V. Parag, Ioannis Lestas
2024 conf
CDC
Shang Wang, Xiaodong Cheng, Peter van Heijster
2024 conf
ECC
Yangming Dou, Xiaodong Cheng, Jacquelien M. A. Scherpen
2024 J jnl
CoRR
Yangming Dou, Xiaodong Cheng, Jacquelien M. A. Scherpen
2023 J jnl
IEEE Trans. Autom. Control.
Xiaodong Cheng, Shengling Shi, Ioannis Lestas, Paul M. J. Van den Hof
2023 J jnl
Autom.
Muhammad Umar B. Niazi, Xiaodong Cheng, Carlos Canudas-de-Wit, Jacquelien M. A. Scherpen
2023 conf
ICIEAI
Haifeng Wang, Liang Zhao, Xiaodong Cheng
2023 J jnl
J. Real Time Image Process.
Jiaocheng Ma, Xiaodong Cheng
2023 J jnl
IEEE Trans. Autom. Control.
Xiaodong Cheng, Lanlin Yu, Dingchao Ren, Jacquelien M. A. Scherpen
2023 J jnl
IEEE Trans. Autom. Control.
Shengling Shi, Xiaodong Cheng, Paul M. J. Van den Hof
2022 J jnl
IEEE Trans. Autom. Control.
Xiaodong Cheng, Shengling Shi, Paul M. J. Van den Hof
2022 J jnl
IEEE Control. Syst. Lett.
H. J. Dreef, Shengling Shi, Xiaodong Cheng, M. C. F. Donkers, Paul M. J. Van den Hof
2022 J jnl
CoRR
H. J. Dreef, Shengling Shi, Xiaodong Cheng, M. C. F. Donkers, Paul M. J. Van den Hof
2022 J jnl
IEEE Trans. Cybern.
Ning Zhou, Xiaodong Cheng, Zhongqi Sun, Yuanqing Xia
2022 J jnl
IEEE Trans. Cybern.
Ning Zhou, Xiaodong Cheng, Yuanqing Xia, Yan-Jun Liu
2022 J jnl
Autom.
Shengling Shi, Xiaodong Cheng, Paul M. J. Van den Hof
2022 J jnl
Autom.
Lanlin Yu, Xiaodong Cheng, Jacquelien M. A. Scherpen, Junlin Xiong
2022 conf
CAIBDA
Chenjie Su, Xiaodong Cheng, Shiqi Xi, Bomeng Li
2021 J jnl
Annu. Rev. Control. Robotics Auton. Syst.
Xiaodong Cheng, Jacquelien M. A. Scherpen
2020 conf
RICAI
Xi Wang, Xiaodong Cheng, ZheFu Chen, Fei Xu
2020 conf
ECC
Xiaodong Cheng, Ion Necoara
2020 J jnl
IEEE Trans. Autom. Control.
Xiaodong Cheng, Jacquelien M. A. Scherpen
2020 conf
CDC
Ning Zhou, Xiaodong Cheng, Yuanqing Xia, Yanjun Liu
2020 J jnl
CoRR
Shengling Shi, Xiaodong Cheng, Paul M. J. Van den Hof
2020 J jnl
Digit. Signal Process.
Yang Liu, Yuting Li, Xiaodong Cheng, Yinbo Lian, Yongjun Jia, Hui Zhang
2020 J jnl
Autom.
Xiaodong Cheng, Jacquelien M. A. Scherpen
2020 J jnl
CoRR
Shengling Shi, Xiaodong Cheng, Paul M. J. Van den Hof
2020 conf
RICAI
Shuai Hou, Xiaodong Cheng, Lei Shi, Shiqiang Zhang
2019 J jnl
IEEE Trans. Control. Syst. Technol.
Michele Cucuzzella, Sebastian Trip, Claudio De Persis, Xiaodong Cheng, Antonella Ferrara, Arjan van der Schaft
2019 conf
CDC
Xiaodong Cheng, Shengling Shi, Paul M. J. Van den Hof
2019 J jnl
Appl. Soft Comput.
Yonghong Huang, Huan Zang, Xiaodong Cheng, Hongsheng Wu, Jueyou Li
2019 J jnl
IEEE Control. Syst. Lett.
Sebastian Trip, Michele Cucuzzella, Xiaodong Cheng, Jacquelien M. A. Scherpen
2019 J jnl
IEEE Trans. Autom. Control.
Xiaodong Cheng, Yu Kawano, Jacquelien M. A. Scherpen
2019 J jnl
Eur. J. Control
Xiaodong Cheng, Jacquelien M. A. Scherpen, Fan Zhang
2019 C conf
ACC
Liangming Chen, Ming Cao, Chuanjiang Li, Xiaodong Cheng, Yuri A. Kapitanyuk
2019 conf
ECC
Xiaodong Cheng, Lanlin Yu, Jacquelien M. A. Scherpen
2019 conf
CDC
Muhammad Umar B. Niazi, Xiaodong Cheng, Carlos Canudas-de-Wit, Jacquelien M. A. Scherpen
2019 conf
CDC
Lanlin Yu, Xiaodong Cheng, Jacquelien M. A. Scherpen, Junlin Xiong
2019 conf
CDC
Lanlin Yu, Xiaodong Cheng, Jacquelien M. A. Scherpen, Emma Gort
2018 J jnl
IEEE Access
Heng Wang, Zhenzhen Zhao, Xiaodong Cheng, Jilai Ying, Jianhua Qu, Guangyin Xu
2018 J jnl
Adv. Comput. Math.
Xiaodong Cheng, Jacquelien M. A. Scherpen
2018 conf
ICSAI
Ning Ding, Xiaodong Cheng, Zhengyi Cui
2018 conf
CSAE
Dongsheng Li, Xiaodong Cheng, Mengsongqi Li
2018 conf
ICCIP
Zhengyi Cui, Xiaodong Cheng, Ning Ding, Wenhua Wang, Xiao-Fang Wang, Mengsongqi Li
2018 conf
ECC
Xiaodong Cheng, Jacquelien M. A. Scherpen
2018 C conf
ACC
Sebastian Trip, Michele Cucuzzella, Claudio De Persis, Xiaodong Cheng, Antonella Ferrara
2018 conf
CSAE
Zhen Yang, Xiaodong Cheng, Mengsongqi Li
2017 conf
CDC
Xiaodong Cheng, Jacquelien M. A. Scherpen
2017 J jnl
IEEE Trans. Autom. Control.
Xiaodong Cheng, Yu Kawano, Jacquelien M. A. Scherpen
2017 J jnl
CoRR
Xiaodong Cheng, Yu Kawano, Jacquelien M. A. Scherpen
2016 conf
ECC
Xiaodong Cheng, Yu Kawano, Jacquelien M. A. Scherpen
2016 conf
CDC
Xiaodong Cheng, Jacquelien M. A. Scherpen
2016 conf
CDC
Xiaodong Cheng, Jacquelien M. A. Scherpen, Yu Kawano
2013 J jnl
Int. J. Web Appl.
Xiping Zhao, Xiaodong Cheng, Xiaofei Li
2012 conf
FSKD
Xiaodong Cheng, Xiaojing Sun, Zhiyong Li
2012 C conf
INDIN
Yushou Gai, Yan Shi, Maolin Cai, Xiangheng Fu, Xiaodong Cheng
1994 J jnl
Nucleic Acids Res.
Sanjay Kumar, Xiaodong Cheng, Saulius Klimasauskas, Sha Mi, Janos Posfai, Richard J. Roberts, Geoffrey G. Wilson
redb/extractors/js_extractors/js_context.py
← Index redb/extractors/js_extractors/js_context.py python
"""Per-sample shared state for the JavaScript extractor pipeline.

A `JSContext` is built exactly once per JS sample (in `workers.py`) and threaded
into every extractor that runs against that sample. It owns the disk read, the
decoded source text, the line-split cache, the Shannon text-entropy figure, the
shared `scan_source()` results, and the pyjsparser AST. Each of those is
computed lazily through `cached_property` so an extractor that doesn't need a
particular artefact does not pay for it.

Without this object, every JS extractor instance redoes the same disk read,
decode, scan, and (for any consumer) AST parse. With it, every extractor
shares one set of results.

`JSExtractor.__init__` accepts the context via a `context=` kwarg; if absent
(e.g. unit tests instantiating an extractor directly with `source=...`) it
builds a fresh context from the constructor arguments. Either path produces a
fully-populated context, so extractor code can always rely on
`self._context.scan` / `self._context.ast` / etc.
"""

from __future__ import annotations

import math
from collections import Counter
from dataclasses import dataclass
from functools import cached_property
from typing import Any, Dict, List, Optional

import chardet

from redb.extractors.js_extractors.js_patterns import scan_source


def decode_source(raw_bytes: bytes) -> str:
    """Decode raw JS bytes to text, honouring BOMs and falling back to chardet.

    Mirrors the historical `JSExtractor._decode_source` logic so existing tests
    continue to round-trip identically.
    """
    if not raw_bytes:
        return ""

    if raw_bytes[:3] == b"\xef\xbb\xbf":
        return raw_bytes[3:].decode("utf-8", errors="replace")
    if raw_bytes[:2] in (b"\xff\xfe", b"\xfe\xff"):
        return raw_bytes.decode("utf-16", errors="replace")

    try:
        return raw_bytes.decode("utf-8")
    except UnicodeDecodeError:
        pass

    try:
        detected = chardet.detect(raw_bytes)
        if detected and detected.get("encoding"):
            return raw_bytes.decode(detected["encoding"], errors="replace")
    except Exception:
        pass

    return raw_bytes.decode("latin-1", errors="replace")


def _text_entropy(text: str) -> float:
    """Shannon entropy of the character distribution of `text`, rounded to 4dp."""
    if not text:
        return 0.0
    counter = Counter(text)
    length = len(text)
    entropy = 0.0
    for count in counter.values():
        p = count / length
        if p > 0:
            entropy -= p * math.log2(p)
    return round(entropy, 4)


@dataclass
class JSContext:
    """Shared raw materials for one JS sample, consumed by every JS extractor.

    Cheap attributes (raw_bytes, source) are populated eagerly by the factory.
    Expensive ones (scan, ast) are cached_property — computed on first access
    and reused across every extractor that holds the same context.

    `content_type` is the magika label (e.g. `"javascript"`) carried alongside
    the source so the new code_text_content writer (and any future generic
    text-content writer) can record it without re-running magika. Defaults to
    `"javascript"` because by construction this context type is JS-specific;
    workers.py supplies the actual magika value when it builds the context.
    """

    filepath: str
    raw_bytes: bytes
    source: str
    log: Any = None
    content_type: str = "javascript"
    # Populated by JSStringsExtractor.extract() (the decoded/reconstructed
    # strings — hex/unicode/charcode/base64/concat unpacked into plaintext).
    # Read post-loop by the IOC plumbing in workers.py so any IOCs hidden
    # behind those encodings get scraped from the decoded form. Stays None
    # if JSStringsExtractor didn't run for this sample.
    decoded_strings: Optional[list] = None

    @cached_property
    def lines(self) -> List[str]:
        return self.source.splitlines() if self.source else []

    @cached_property
    def text_entropy(self) -> float:
        return _text_entropy(self.source)

    @cached_property
    def scan(self) -> Dict[str, Dict[str, object]]:
        """Result of running scan_source() exactly once over self.source."""
        return scan_source(self.source) if self.source else {}

    @cached_property
    def ast(self) -> Optional[Any]:
        """Lazy pyjsparser AST. Returns None if the parser is missing or fails.

        Extractors should treat None AST as "fall back to regex" — every
        AST-consuming extractor already handles that path.
        """
        if not self.source:
            return None
        try:
            import pyjsparser
            return pyjsparser.parse(self.source)
        except ImportError:
            if self.log is not None:
                self.log.debug("pyjsparser not installed, AST analysis skipped")
        except Exception as e:
            if self.log is not None:
                self.log.warning(f"AST parsing failed for {self.filepath}: {e}")
        return None

    @cached_property
    def deobfuscated(self) -> "tuple[Optional[str], Optional[str]]":
        """Run the configured JS deobfuscator (with jsbeautifier fallback) once
        per sample and cache the result. Returns `(text, normalizer_used)` or
        `(None, None)` if neither path produced output.

        Computed lazily on first access — samples whose pipeline never reads
        this don't pay the subprocess cost.
        """
        from redb.extractors.js_extractors.js_deobfuscator import deobfuscate
        return deobfuscate(self.source, self.log)

    @cached_property
    def scan_deobfuscated(self) -> Dict[str, Dict[str, object]]:
        """Result of running scan_source() exactly once over the deobfuscated
        text, keyed by PATTERNS only (FEATURE_PATTERNS are not consulted by
        the dual-pass consumers). Empty dict when there is no deobfuscated
        text or it equals the raw source.

        Two extractors consume the post-deobf API surface:
        `JSSuspiciousAPIsExtractor` (for revealed_by_deobf rows) and
        `JSDeobfuscationExtractor` (for the new_apis_found diff). Caching here
        means we scan the deobfuscated text once instead of twice per sample.
        """
        from redb.extractors.js_extractors.js_patterns import PATTERNS
        deobf_text, _ = self.deobfuscated
        if not deobf_text or deobf_text == self.source:
            return {}
        return scan_source(deobf_text, patterns=(PATTERNS,))

    @cached_property
    def xray(self):
        """Run @nodesecure/js-x-ray once per sample and cache the result.

        Returns an `XRayResult` (always — the function collapses every failure
        path to an empty result so callers don't have to special-case missing
        Node, missing package, timeouts, or parse errors). The
        `JSFeaturesExtractor` reads it for the obfuscator family name and for
        corroborating warning kinds; the heuristic falls back cleanly when
        `xray.obfuscator is None`.
        """
        from redb.extractors.js_extractors.js_xray import run
        return run(self.source, self.log)

    @classmethod
    def from_path(
        cls,
        filepath: str,
        log: Any = None,
        source: Optional[str] = None,
        raw_bytes: Optional[bytes] = None,
        content_type: str = "javascript",
    ) -> "JSContext":
        """Build a context from disk. `raw_bytes` and `source` are optional
        overrides — useful when the caller has already read or decoded the file.
        `content_type` is the magika label workers.py dispatched on; it lands
        on the context for the code_text_content writer to record.
        """
        if raw_bytes is None:
            with open(filepath, "rb") as f:
                raw_bytes = f.read()
        if source is None:
            source = decode_source(raw_bytes)
        return cls(
            filepath=filepath,
            raw_bytes=raw_bytes,
            source=source,
            log=log,
            content_type=content_type,
        )