Naveed Khan

53 papers B 1C 4Journal 35Unranked 12
YearRankTypeTitle / Venue / Authors
2026 J jnl
Empir. Softw. Eng.
Radowanul Haque, Aftab Ali, Sally I. McClean, Naveed Khan
2026 J jnl
IEEE Wirel. Commun.
Shumaila Javaid, Naveed Khan, Abdulmalik Alwarafy, Nasir Saeed
2026 J jnl
IEEE Trans. Cloud Comput.
Naveed Khan, Junchao Ma, Dingyuan Tang, Mustafa A. Al Sibahee, Shehzad Ashraf Chaudhry, Ashok Kumar Das
2025 J jnl
Discov. Artif. Intell.
Wasan Al Kishri, Jabar H. Yousif, Mahmood Al-Bahri, Muhammad Zakarya, Naveed Khan, Sanad Sulaiman Al Maskari, Ahmet Gurhanli
2025 J jnl
Int. J. Crit. Infrastructure Prot.
Umar Islam, Hanif Ullah, Naveed Khan, Kashif Saleem, Iftikhar Ahmad
2025 J jnl
Qual. Reliab. Eng. Int.
Babar Zaman, Naveed Khan
2025 J jnl
IEEE Trans. Consumer Electron.
Umar Islam, Hanif Ullah, Naveed Khan, Iftikhar Ahmad, Kashif Saleem
2025 J jnl
Image Vis. Comput.
Umar Islam, Hathal Salamah Alwageed, Saleh Alyahyan, Manal Alghieth, Hanif Ullah, Naveed Khan
2025 conf
MeditCom
Naveed Khan, Abdulmalik Alwarafy, Mohamed Hayajneh, Farag Sallabi, Hesham El-Sayed
2025 J jnl
Internet Things
Umar Islam, Mohammed Naif Alatawi, Sulaiman Alamro, Hathal Salamah Alwageed, Hanif Ullah, Naveed Khan
2025 J jnl
CoRR
Radowanul Haque, Aftab Ali, Sally I. McClean, Naveed Khan
2025 J jnl
Int. J. Comput. Intell. Syst.
Falah Y. H. Ahmed, Muhammad Zakarya, Naveed Khan, Dilovan Asaad Zebari, Mahmood Al-Bahri, Bwalya Kelvin Joseph, Abdullah Abdullah
2025 J jnl
IEEE Access
Abdulmalik Alwarafy, Naveed Khan, Nasir Saeed, Mustaqeem Khan, Zahid Mehmood, Saeed Abdallah
2025 J jnl
IEEE Access
Naveed Khan, Kamran Mujahid, Syed Ali Abbas Kazmi, Jawad Khan, Fahad Alturise
2024 conf
UCAmI
Majid Liaquat, Chris D. Nugent, Ian Cleland, Naveed Khan
2024 J jnl
IEEE Internet Things J.
Mustafa A. Al Sibahee, Zaid Ameen Abduljabbar, Alladoumbaye Ngueilbaye, Chengwen Luo, Jianqiang Li, Yijing Huang, Jin Zhang, Naveed Khan, Vincent Omollo Nyangaresi, Ali Hasan Ali
2024 J jnl
Qual. Reliab. Eng. Int.
Babar Zaman, Hafiz Zafar Nazir, Naveed Khan, Muhammad Riaz, Tahir Abbas
2024 J jnl
IEEE Access
Naveed Khan, Ayaz Ahmad, Abdul Wakeel, Zeeshan Kaleem, Bushra Rashid, Waqas Khalid
2024 J jnl
CoRR
Naveed Khan, Ayaz Ahmad, Abdul Wakeel, Zeeshan Kaleem, Bushra Rashid, Waqas Khalid
2024 J jnl
Eng. Appl. Artif. Intell.
Yongyu Dai, Zhengwei Huang, Weijun He, Naveed Khan, Yang Yang
2024 J jnl
Telecommun. Syst.
Naveed Khan, Ayaz Ahmad, Sher Ali, Tariqullah Jan, Ihsan Ullah
2024 J jnl
Expert Syst. Appl.
Kamil Shah, Wenqi Liu, Aeshah A. Raezah, Naveed Khan, Sami Ullah Khan, Muhammad Ozair, Zubair Ahmad
2023 J jnl
Comput. Ind. Eng.
Rashid Mehmood, Kassimu Mpungu, Iftikhar Ali, Babar Zaman, Fawwad Hussain Qureshi, Naveed Khan
2023 J jnl
J. Cloud Comput.
Naveed Khan, Jianbiao Zhang, Huhnkuk Lim, Jehad Ali, Intikhab Ullah, Muhammad Salman Pathan, Shehzad Ashraf Chaudhry
2023 J jnl
Qual. Reliab. Eng. Int.
Babar Zaman, Syed Zeeshan Mahfooz, Rashid Mehmood, Naveed Khan, Tajammal Imran
2023 J jnl
Comput. Ind. Eng.
Babar Zaman, Syed Zeeshan Mahfooz, Naveed Khan, Muhammad Riaz, Rashid Mehmood
2023 J jnl
Bus. Process. Manag. J.
Sajjad Alam, Jianhua Zhang, Muhammad Usman Shehzad, Naveed Khan, Ahmad Ali
2023 J jnl
J. Cloud Comput.
Naveed Khan, Sally I. McClean, Shuai Zhang, Chris D. Nugent
2021 J jnl
Autom. Softw. Eng.
Aftab Ali, Naveed Khan, Mamun I. Abu-Tair, Joost Noppen, Sally I. McClean, Ian R. McChesney
2021 conf
RTIP2R
Naveed Khan, Zeeshan Tariq, Aftab Ali, Sally I. McClean, Paul N. Taylor, Detlef D. Nauck
2021 J jnl
Multim. Tools Appl.
Irfan Ahmed, Aftab Khan, Asfandyar Khan, Kamran Mujahid, Naveed Khan
2020 J jnl
Algorithms
Zeeshan Tariq, Naveed Khan, Darryl Charles, Sally I. McClean, Ian R. McChesney, Paul N. Taylor
2019 conf
SGAI Conf.
Naveed Khan, Zulfiqar Ali, Aftab Ali, Sally I. McClean, Darryl Charles, Paul N. Taylor, Detlef D. Nauck
2019 C conf
CLOSER
Naveed Khan, Raju Shrestha
2019 C conf
MEDES
Naveed Khan, Hårek Haugerud, Raju Shrestha, Anis Yazidi
2019 J jnl
Future Gener. Comput. Syst.
Zulfiqar Ali, Muhammad Imran, Sally I. McClean, Naveed Khan, Muhammad Shoaib
2019 C conf
ICORES
Sally I. McClean, David A. Stanford, Lalit Garg, Naveed Khan
2018 conf
SGAI Conf.
Sally I. McClean, Naveed Khan, Adam R. Currie, Kashaf Khan
2018
Naveed Khan
2018 J jnl
Mob. Inf. Syst.
Moneeb Gohar, Naveed Khan, Awais Ahmad, Muhammad Najam-ul-Islam, Shahzad Sarwar, Seok Joo Koh
2018 J jnl
Clust. Comput.
Saleh M. Al-Saleem, Aftab Ali, Naveed Khan
2017 J jnl
IEEE Trans. Mob. Comput.
Timothy Patterson, Naveed Khan, Sally I. McClean, Chris D. Nugent, Shuai Zhang, Ian Cleland, Qin Ni
2016 J jnl
Mob. Networks Appl.
Udsanee Pakdeetrakulwong, Pornpit Wongthongtham, Waralak V. Siricharoen, Naveed Khan
2016 conf
UCAmI (1)
Naveed Khan, Sally I. McClean, Shuai Zhang, Chris D. Nugent
2016 J jnl
Sensors
Naveed Khan, Sally I. McClean, Shuai Zhang, Chris D. Nugent
2016 B conf
CBMS
Naveed Khan, Sally I. McClean, Shuai Zhang, Chris D. Nugent
2015 conf
ASWEC (2)
Udsanee Pakdeetrakulwong, Pornpit Wongthongtham, Naveed Khan
2015 conf
MWSCAS
Naveed Khan, Hesham Omran, Yingbang Yao, Khaled N. Salama
2015 conf
UCAmI
Naveed Khan, Sally I. McClean, Shuai Zhang, Chris D. Nugent
2013 C conf
SNPD
Abdulaziz Alsadhan, Naveed Khan
2013 conf
ICNSC
Ihsan Ullah, Naveed Khan, Hatim A. Aboalsamh
2010 conf
FGIT
Maqsood Mahmud, Abdulrahman Abdulkarim Mirza, Ihsan Ullah, Naveed Khan, Abdul Hanan Bin Abdullah, Mohammad Yazid Bin Idris
2009 conf
CICC
Amir Bashir, Jing Li, Kiran Ivatury, Naveed Khan, Nirav Gala, Noam Familia, Zulfiqar Mohammed
redb/extractors/js_extractors/js_context.py
← Index redb/extractors/js_extractors/js_context.py python
"""Per-sample shared state for the JavaScript extractor pipeline.

A `JSContext` is built exactly once per JS sample (in `workers.py`) and threaded
into every extractor that runs against that sample. It owns the disk read, the
decoded source text, the line-split cache, the Shannon text-entropy figure, the
shared `scan_source()` results, and the pyjsparser AST. Each of those is
computed lazily through `cached_property` so an extractor that doesn't need a
particular artefact does not pay for it.

Without this object, every JS extractor instance redoes the same disk read,
decode, scan, and (for any consumer) AST parse. With it, every extractor
shares one set of results.

`JSExtractor.__init__` accepts the context via a `context=` kwarg; if absent
(e.g. unit tests instantiating an extractor directly with `source=...`) it
builds a fresh context from the constructor arguments. Either path produces a
fully-populated context, so extractor code can always rely on
`self._context.scan` / `self._context.ast` / etc.
"""

from __future__ import annotations

import math
from collections import Counter
from dataclasses import dataclass
from functools import cached_property
from typing import Any, Dict, List, Optional

import chardet

from redb.extractors.js_extractors.js_patterns import scan_source


def decode_source(raw_bytes: bytes) -> str:
    """Decode raw JS bytes to text, honouring BOMs and falling back to chardet.

    Mirrors the historical `JSExtractor._decode_source` logic so existing tests
    continue to round-trip identically.
    """
    if not raw_bytes:
        return ""

    if raw_bytes[:3] == b"\xef\xbb\xbf":
        return raw_bytes[3:].decode("utf-8", errors="replace")
    if raw_bytes[:2] in (b"\xff\xfe", b"\xfe\xff"):
        return raw_bytes.decode("utf-16", errors="replace")

    try:
        return raw_bytes.decode("utf-8")
    except UnicodeDecodeError:
        pass

    try:
        detected = chardet.detect(raw_bytes)
        if detected and detected.get("encoding"):
            return raw_bytes.decode(detected["encoding"], errors="replace")
    except Exception:
        pass

    return raw_bytes.decode("latin-1", errors="replace")


def _text_entropy(text: str) -> float:
    """Shannon entropy of the character distribution of `text`, rounded to 4dp."""
    if not text:
        return 0.0
    counter = Counter(text)
    length = len(text)
    entropy = 0.0
    for count in counter.values():
        p = count / length
        if p > 0:
            entropy -= p * math.log2(p)
    return round(entropy, 4)


@dataclass
class JSContext:
    """Shared raw materials for one JS sample, consumed by every JS extractor.

    Cheap attributes (raw_bytes, source) are populated eagerly by the factory.
    Expensive ones (scan, ast) are cached_property — computed on first access
    and reused across every extractor that holds the same context.

    `content_type` is the magika label (e.g. `"javascript"`) carried alongside
    the source so the new code_text_content writer (and any future generic
    text-content writer) can record it without re-running magika. Defaults to
    `"javascript"` because by construction this context type is JS-specific;
    workers.py supplies the actual magika value when it builds the context.
    """

    filepath: str
    raw_bytes: bytes
    source: str
    log: Any = None
    content_type: str = "javascript"
    # Populated by JSStringsExtractor.extract() (the decoded/reconstructed
    # strings — hex/unicode/charcode/base64/concat unpacked into plaintext).
    # Read post-loop by the IOC plumbing in workers.py so any IOCs hidden
    # behind those encodings get scraped from the decoded form. Stays None
    # if JSStringsExtractor didn't run for this sample.
    decoded_strings: Optional[list] = None

    @cached_property
    def lines(self) -> List[str]:
        return self.source.splitlines() if self.source else []

    @cached_property
    def text_entropy(self) -> float:
        return _text_entropy(self.source)

    @cached_property
    def scan(self) -> Dict[str, Dict[str, object]]:
        """Result of running scan_source() exactly once over self.source."""
        return scan_source(self.source) if self.source else {}

    @cached_property
    def ast(self) -> Optional[Any]:
        """Lazy pyjsparser AST. Returns None if the parser is missing or fails.

        Extractors should treat None AST as "fall back to regex" — every
        AST-consuming extractor already handles that path.
        """
        if not self.source:
            return None
        try:
            import pyjsparser
            return pyjsparser.parse(self.source)
        except ImportError:
            if self.log is not None:
                self.log.debug("pyjsparser not installed, AST analysis skipped")
        except Exception as e:
            if self.log is not None:
                self.log.warning(f"AST parsing failed for {self.filepath}: {e}")
        return None

    @cached_property
    def deobfuscated(self) -> "tuple[Optional[str], Optional[str]]":
        """Run the configured JS deobfuscator (with jsbeautifier fallback) once
        per sample and cache the result. Returns `(text, normalizer_used)` or
        `(None, None)` if neither path produced output.

        Computed lazily on first access — samples whose pipeline never reads
        this don't pay the subprocess cost.
        """
        from redb.extractors.js_extractors.js_deobfuscator import deobfuscate
        return deobfuscate(self.source, self.log)

    @cached_property
    def scan_deobfuscated(self) -> Dict[str, Dict[str, object]]:
        """Result of running scan_source() exactly once over the deobfuscated
        text, keyed by PATTERNS only (FEATURE_PATTERNS are not consulted by
        the dual-pass consumers). Empty dict when there is no deobfuscated
        text or it equals the raw source.

        Two extractors consume the post-deobf API surface:
        `JSSuspiciousAPIsExtractor` (for revealed_by_deobf rows) and
        `JSDeobfuscationExtractor` (for the new_apis_found diff). Caching here
        means we scan the deobfuscated text once instead of twice per sample.
        """
        from redb.extractors.js_extractors.js_patterns import PATTERNS
        deobf_text, _ = self.deobfuscated
        if not deobf_text or deobf_text == self.source:
            return {}
        return scan_source(deobf_text, patterns=(PATTERNS,))

    @cached_property
    def xray(self):
        """Run @nodesecure/js-x-ray once per sample and cache the result.

        Returns an `XRayResult` (always — the function collapses every failure
        path to an empty result so callers don't have to special-case missing
        Node, missing package, timeouts, or parse errors). The
        `JSFeaturesExtractor` reads it for the obfuscator family name and for
        corroborating warning kinds; the heuristic falls back cleanly when
        `xray.obfuscator is None`.
        """
        from redb.extractors.js_extractors.js_xray import run
        return run(self.source, self.log)

    @classmethod
    def from_path(
        cls,
        filepath: str,
        log: Any = None,
        source: Optional[str] = None,
        raw_bytes: Optional[bytes] = None,
        content_type: str = "javascript",
    ) -> "JSContext":
        """Build a context from disk. `raw_bytes` and `source` are optional
        overrides — useful when the caller has already read or decoded the file.
        `content_type` is the magika label workers.py dispatched on; it lands
        on the context for the code_text_content writer to record.
        """
        if raw_bytes is None:
            with open(filepath, "rb") as f:
                raw_bytes = f.read()
        if source is None:
            source = decode_source(raw_bytes)
        return cls(
            filepath=filepath,
            raw_bytes=raw_bytes,
            source=source,
            log=log,
            content_type=content_type,
        )