Xiande Zhao

61 papers Journal 58Unranked 3
YearRankTypeTitle / Venue / Authors
2025 J jnl
Comput. Electron. Agric.
Enhui Wu, Yu Chen, Ruijun Ma, Xiande Zhao
2025 J jnl
Comput. Electron. Agric.
Zhen Gao, Daming Dong, Guiyan Yang, Xuelin Wen, Juekun Bai, Fengjing Cao, Chunjiang Zhao, Xiande Zhao
2025 J jnl
Ind. Manag. Data Syst.
Xueyuan Liu, Ying Kei Tse, Yan Yu, Haoliang Huang, Xiande Zhao
2024 J jnl
Eur. J. Oper. Res.
Jinyu Yang, Wenqing Zhang, Xiande Zhao
2024 J jnl
Comput. Electron. Agric.
Zhen Gao, Guiyan Yang, Xiande Zhao, Leizi Jiao, Xuelin Wen, Yachao Liu, Xintao Xia, Chunjiang Zhao, Daming Dong
2024 J jnl
IEEE Trans. Engineering Management
Yina Li, Yang Tong, Fei Ye, Ying-Ju Chen, Xiande Zhao
2022 J jnl
Ann. Oper. Res.
Xiang Li, Tianyu Zhang, Liang Wang, Hongguang Ma, Xiande Zhao
2022 J jnl
Ind. Manag. Data Syst.
Qian Yang, Liping Qian, Xiande Zhao
2022 J jnl
Int. J. Prod. Res.
Yi Zhang, Xiang Li, Liang Wang, Xiande Zhao, Jinwu Gao
2022 J jnl
IEEE Trans. Engineering Management
Zhiqiang Wang, Tobias Schoenherr, Xiande Zhao, Shanshan Zhang
2022 J jnl
Ind. Manag. Data Syst.
Siyu Li, Kedi Wang, Baofeng Huo, Xiande Zhao, Xiling Cui
2021 J jnl
Ind. Manag. Data Syst.
Hao Ying, Lujie Chen, Xiande Zhao
2021 J jnl
Ind. Manag. Data Syst.
Jiandong Zhou, Xiang Li, Xiande Zhao, Liang Wang
2020 J jnl
Ind. Manag. Data Syst.
Haidi Zhou, Qiang Wang, Xiande Zhao
2020 J jnl
Ind. Manag. Data Syst.
Xiao Song, Hao Ying, Xiande Zhao, Lujie Chen
2020 J jnl
Ind. Manag. Data Syst.
Qian Yang, Liping Qian, Xiande Zhao
2020 J jnl
IEEE Trans. Engineering Management
Zhiqiang Wang, Ruey-Jer "Bryan" Jean, Xiande Zhao
2020 J jnl
Ind. Manag. Data Syst.
Qiuping Huang, Xiande Zhao, Min Zhang, KwanHo Yeung, Lijun Ma, Jeff Hoi Yan Yeung
2019 J jnl
IEEE Trans. Engineering Management
Qiang Wang, Chris Voss, Xiande Zhao
2019 J jnl
Ind. Manag. Data Syst.
Siyu Li, Xiling Cui, Baofeng Huo, Xiande Zhao
2019 J jnl
Ind. Manag. Data Syst.
Yanming Zhang, Xiande Zhao, Baofeng Huo
2018 J jnl
Ind. Manag. Data Syst.
Xueyuan Liu, Haiyun Zhao, Xiande Zhao
2018 J jnl
Int. J. Serv. Technol. Manag.
Haiju Hu, Xiande Zhao
2018 J jnl
Ind. Manag. Data Syst.
Wenhui Fu, Qiang Wang, Xiande Zhao
2018 J jnl
Ind. Manag. Data Syst.
Wenhui Fu, Qiang Wang, Xiande Zhao
2018 J jnl
Ind. Manag. Data Syst.
Min Li, Zhiqiang Wang, Xiande Zhao
2017 J jnl
Ind. Manag. Data Syst.
Qianling Chen, Min Zhang, Xiande Zhao
2017 J jnl
Ind. Manag. Data Syst.
Baofeng Huo, Qianwen Wang, Xiande Zhao, Zhongsheng Hua
2017 conf
CCTA (1)
Yinchi Ma, Wen Ding, Yonghua Qu, Xiande Zhao
2017 J jnl
Ind. Manag. Data Syst.
Haiju Hu, Ramdane Djebarni, Xiande Zhao, Liwei Xiao, Barbara B. Flynn
2017 J jnl
Ind. Manag. Data Syst.
Shanshan Zhang, Zhiqiang Wang, Xiande Zhao, Min Zhang
2017 J jnl
Int. J. Prod. Res.
Min Zhang, Hangfei Guo, Xiande Zhao
2016 J jnl
Ind. Manag. Data Syst.
Zhiqiang Wang, Baofeng Huo, Yinan Qi, Xiande Zhao
2016 J jnl
Comput. Electron. Agric.
Pengcheng Han, Daming Dong, Xiande Zhao, Leizi Jiao, Yun Lang
2016 J jnl
Ind. Manag. Data Syst.
Zhiqiang Wang, Qiang Wang, Xiande Zhao, Marjorie A. Lyles, Guilong Zhu
2016 J jnl
Comput. Electron. Agric.
Leizi Jiao, Daming Dong, Haikuan Feng, Xiande Zhao, Liping Chen
2015 conf
CCTA (1)
Fei Liu, Peichao Zheng, Baichuan Huang, Xiande Zhao, Leizi Jiao, Daming Dong
2015 J jnl
Ind. Manag. Data Syst.
Zhi Cao, Baofeng Huo, Yuan Li, Xiande Zhao
2015 J jnl
Ind. Manag. Data Syst.
Chen Liu, Baofeng Huo, Shulin Liu, Xiande Zhao
2015 J jnl
Ind. Manag. Data Syst.
Xiande Zhao, KwanHo Yeung, Qiuping Huang, Xiao Song
2015 J jnl
Ind. Manag. Data Syst.
Qiang Wang, Chris Voss, Xiande Zhao, Zhiqiang Wang
2015 J jnl
Inf. Manag.
Baofeng Huo, Cheng Zhang, Xiande Zhao
2014 J jnl
IEEE Trans. Engineering Management
Baofeng Huo, Xiande Zhao, Fujun Lai
2012 J jnl
Eur. J. Oper. Res.
Yina Li, Xuejun Xu, Xiande Zhao, Jeff Hoi Yan Yeung, Fei Ye
2012 J jnl
IEEE Trans. Engineering Management
Fujun Lai, Min Zhang, Denis M. S. Lee, Xiande Zhao
2011 J jnl
Decis. Sci.
Yinan Qi, Xiande Zhao, Chwen Sheu
2010 J jnl
Eur. J. Oper. Res.
Jinxing Xie, Deming Zhou, Jerry C. Wei, Xiande Zhao
2009 J jnl
Decis. Sci.
Yinan Qi, Kenneth K. Boyer, Xiande Zhao
2008 J jnl
Comput. Ind. Eng.
R. S. M. Lau, Jinxing Xie, Xiande Zhao
2008 J jnl
Decis. Sci.
Manus Johnny Rungtusanatham, C. H. Ng, Xiande Zhao, T. S. Lee
2007 J jnl
Int. J. Uncertain. Fuzziness Knowl. Based Syst.
Xiaoyu Ji, Xiande Zhao, Deming Zhou
2007 J jnl
Decis. Sci.
Xiande Zhao, Barbara B. Flynn, Aleda V. Roth
2007 J jnl
Int. J. Uncertain. Fuzziness Knowl. Based Syst.
Baoding Liu, Xiande Zhao
2007 J jnl
Eur. J. Oper. Res.
Xiande Zhao, Jinxing Xie, Jerry C. Wei
2006 J jnl
Decis. Sci.
Xiande Zhao, Barbara B. Flynn, Aleda V. Roth
2006 J jnl
Int. J. Internet Enterp. Manag.
Fujun Lai, Xiande Zhao, Tien-Sheng Lee
2006 J jnl
Ind. Manag. Data Syst.
Fujun Lai, Xiande Zhao, Qiang Wang
2004 J jnl
Comput. Ind. Eng.
Jinxing Xie, T. S. Lee, Xiande Zhao
2004 conf
ICEB
Baofeng Huo, Xiande Zhao, Jeff Hoi Yan Yeung
2002 J jnl
Decis. Sci.
Xiande Zhao, Jinxing Xie, Jerry C. Wei
2002 J jnl
Eur. J. Oper. Res.
Xiande Zhao, Jinxing Xie, Janny Leung
redb/extractors/js_extractors/js_context.py
← Index redb/extractors/js_extractors/js_context.py python
"""Per-sample shared state for the JavaScript extractor pipeline.

A `JSContext` is built exactly once per JS sample (in `workers.py`) and threaded
into every extractor that runs against that sample. It owns the disk read, the
decoded source text, the line-split cache, the Shannon text-entropy figure, the
shared `scan_source()` results, and the pyjsparser AST. Each of those is
computed lazily through `cached_property` so an extractor that doesn't need a
particular artefact does not pay for it.

Without this object, every JS extractor instance redoes the same disk read,
decode, scan, and (for any consumer) AST parse. With it, every extractor
shares one set of results.

`JSExtractor.__init__` accepts the context via a `context=` kwarg; if absent
(e.g. unit tests instantiating an extractor directly with `source=...`) it
builds a fresh context from the constructor arguments. Either path produces a
fully-populated context, so extractor code can always rely on
`self._context.scan` / `self._context.ast` / etc.
"""

from __future__ import annotations

import math
from collections import Counter
from dataclasses import dataclass
from functools import cached_property
from typing import Any, Dict, List, Optional

import chardet

from redb.extractors.js_extractors.js_patterns import scan_source


def decode_source(raw_bytes: bytes) -> str:
    """Decode raw JS bytes to text, honouring BOMs and falling back to chardet.

    Mirrors the historical `JSExtractor._decode_source` logic so existing tests
    continue to round-trip identically.
    """
    if not raw_bytes:
        return ""

    if raw_bytes[:3] == b"\xef\xbb\xbf":
        return raw_bytes[3:].decode("utf-8", errors="replace")
    if raw_bytes[:2] in (b"\xff\xfe", b"\xfe\xff"):
        return raw_bytes.decode("utf-16", errors="replace")

    try:
        return raw_bytes.decode("utf-8")
    except UnicodeDecodeError:
        pass

    try:
        detected = chardet.detect(raw_bytes)
        if detected and detected.get("encoding"):
            return raw_bytes.decode(detected["encoding"], errors="replace")
    except Exception:
        pass

    return raw_bytes.decode("latin-1", errors="replace")


def _text_entropy(text: str) -> float:
    """Shannon entropy of the character distribution of `text`, rounded to 4dp."""
    if not text:
        return 0.0
    counter = Counter(text)
    length = len(text)
    entropy = 0.0
    for count in counter.values():
        p = count / length
        if p > 0:
            entropy -= p * math.log2(p)
    return round(entropy, 4)


@dataclass
class JSContext:
    """Shared raw materials for one JS sample, consumed by every JS extractor.

    Cheap attributes (raw_bytes, source) are populated eagerly by the factory.
    Expensive ones (scan, ast) are cached_property — computed on first access
    and reused across every extractor that holds the same context.

    `content_type` is the magika label (e.g. `"javascript"`) carried alongside
    the source so the new code_text_content writer (and any future generic
    text-content writer) can record it without re-running magika. Defaults to
    `"javascript"` because by construction this context type is JS-specific;
    workers.py supplies the actual magika value when it builds the context.
    """

    filepath: str
    raw_bytes: bytes
    source: str
    log: Any = None
    content_type: str = "javascript"
    # Populated by JSStringsExtractor.extract() (the decoded/reconstructed
    # strings — hex/unicode/charcode/base64/concat unpacked into plaintext).
    # Read post-loop by the IOC plumbing in workers.py so any IOCs hidden
    # behind those encodings get scraped from the decoded form. Stays None
    # if JSStringsExtractor didn't run for this sample.
    decoded_strings: Optional[list] = None

    @cached_property
    def lines(self) -> List[str]:
        return self.source.splitlines() if self.source else []

    @cached_property
    def text_entropy(self) -> float:
        return _text_entropy(self.source)

    @cached_property
    def scan(self) -> Dict[str, Dict[str, object]]:
        """Result of running scan_source() exactly once over self.source."""
        return scan_source(self.source) if self.source else {}

    @cached_property
    def ast(self) -> Optional[Any]:
        """Lazy pyjsparser AST. Returns None if the parser is missing or fails.

        Extractors should treat None AST as "fall back to regex" — every
        AST-consuming extractor already handles that path.
        """
        if not self.source:
            return None
        try:
            import pyjsparser
            return pyjsparser.parse(self.source)
        except ImportError:
            if self.log is not None:
                self.log.debug("pyjsparser not installed, AST analysis skipped")
        except Exception as e:
            if self.log is not None:
                self.log.warning(f"AST parsing failed for {self.filepath}: {e}")
        return None

    @cached_property
    def deobfuscated(self) -> "tuple[Optional[str], Optional[str]]":
        """Run the configured JS deobfuscator (with jsbeautifier fallback) once
        per sample and cache the result. Returns `(text, normalizer_used)` or
        `(None, None)` if neither path produced output.

        Computed lazily on first access — samples whose pipeline never reads
        this don't pay the subprocess cost.
        """
        from redb.extractors.js_extractors.js_deobfuscator import deobfuscate
        return deobfuscate(self.source, self.log)

    @cached_property
    def scan_deobfuscated(self) -> Dict[str, Dict[str, object]]:
        """Result of running scan_source() exactly once over the deobfuscated
        text, keyed by PATTERNS only (FEATURE_PATTERNS are not consulted by
        the dual-pass consumers). Empty dict when there is no deobfuscated
        text or it equals the raw source.

        Two extractors consume the post-deobf API surface:
        `JSSuspiciousAPIsExtractor` (for revealed_by_deobf rows) and
        `JSDeobfuscationExtractor` (for the new_apis_found diff). Caching here
        means we scan the deobfuscated text once instead of twice per sample.
        """
        from redb.extractors.js_extractors.js_patterns import PATTERNS
        deobf_text, _ = self.deobfuscated
        if not deobf_text or deobf_text == self.source:
            return {}
        return scan_source(deobf_text, patterns=(PATTERNS,))

    @cached_property
    def xray(self):
        """Run @nodesecure/js-x-ray once per sample and cache the result.

        Returns an `XRayResult` (always — the function collapses every failure
        path to an empty result so callers don't have to special-case missing
        Node, missing package, timeouts, or parse errors). The
        `JSFeaturesExtractor` reads it for the obfuscator family name and for
        corroborating warning kinds; the heuristic falls back cleanly when
        `xray.obfuscator is None`.
        """
        from redb.extractors.js_extractors.js_xray import run
        return run(self.source, self.log)

    @classmethod
    def from_path(
        cls,
        filepath: str,
        log: Any = None,
        source: Optional[str] = None,
        raw_bytes: Optional[bytes] = None,
        content_type: str = "javascript",
    ) -> "JSContext":
        """Build a context from disk. `raw_bytes` and `source` are optional
        overrides — useful when the caller has already read or decoded the file.
        `content_type` is the magika label workers.py dispatched on; it lands
        on the context for the code_text_content writer to record.
        """
        if raw_bytes is None:
            with open(filepath, "rb") as f:
                raw_bytes = f.read()
        if source is None:
            source = decode_source(raw_bytes)
        return cls(
            filepath=filepath,
            raw_bytes=raw_bytes,
            source=source,
            log=log,
            content_type=content_type,
        )