Haiying Liu

112 papers A 2B 6C 1Misc 1Journal 63Unranked 38
YearRankTypeTitle / Venue / Authors
2026 J jnl
IEEE Trans. Aerosp. Electron. Syst.
Haiying Liu, Ruihang Liu, Wei Tan, Ruohan Jia, Tiexiang Li
2026 J jnl
Digit. Signal Process.
Dapeng Wan, Lixia Deng, Jinshun Dong, Meiqi Guo, Jianqin Yin, Chenxu Liu, Haiying Liu
2026 J jnl
Expert Syst. Appl.
Jinshun Dong, Lixia Deng, Dapeng Wan, Chenxu Liu, Jianqin Yin, Meiqi Guo, Hongyu Zhang, Shoujun Lin, Haiying Liu, Lida Liu
2026 J jnl
Biomed. Signal Process. Control.
Guoliang Li, Yibin Li, Yuechong Zhang, Qianqian Li, Fuxin Du, Rui Song, Lei Qi, Haiying Liu, Lixia Deng
2026 J jnl
IEEE Trans. Engineering Management
Wenjie Bi, Jingchao Xie, Liqun Wei, Haiying Liu
2025 J jnl
IEEE Trans Autom. Sci. Eng.
Chen Zhao, Changsheng Li, Mingchang Du, Haolei Fan, Jinpeng Diao, Kaifeng Wang, Jun Wang, Weiwei Xia, Xingguang Duan, Haiying Liu, Yuegang Tan, Tianliang Li
2025 J jnl
IEEE Geosci. Remote. Sens. Lett.
Yifeng Du, Haiying Liu, Yang Shengqi, Junmei Guo, Dehao Dong, Jason Gu, Chaoqun Wang, Lida Liu
2025 J jnl
J. Supercomput.
Shengqi Yang, Haiying Liu, Yifeng Du, Tao Sun, Chaoqun Wang, Fei Lin, Lida Liu
2025 J jnl
Comput. Mater. Continua
Bolin Zhang, Haiying Liu, Ruichun Gu
2025 J jnl
J. Imaging
Haiying Liu, Yingchao Li, Shilong Xu, Haoyu Wang, Qiang Fu, Huilin Jiang
2025 J jnl
RAIRO Oper. Res.
Bing Wang, Wenjie Bi, Haiying Liu
2025 J jnl
Digit. Signal Process.
Zenghui Qu, Haiying Liu, Weigang Kong, Jason Gu, Chaoqun Wang, Lixia Deng, Lida Liu, Fei Lin
2025 J jnl
IEEE Signal Process. Lett.
Enyan Li, Lu Yang, Kun Su, Haiying Liu
2025 J jnl
IEEE Trans. Instrum. Meas.
Sun Tao, Yang Shengqi, Haiying Liu, Jason Gu, Lixia Deng, Lida Liu
2025 J jnl
Unmanned Syst.
Qihui Liu, Kuihua Huang, Haiying Liu, Rui Wang, Guangquan Cheng
2024 J jnl
Int. J. Reason. based Intell. Syst.
Haiying Liu
2024 J jnl
Neural Comput. Appl.
Haiying Liu, Zhaoyi He, Rui Wang, Kuihua Huang, Guangquan Cheng
2024 J jnl
IEEE Access
Tao Sun, Haonan Chen, Haiying Liu, Lixia Deng, Lida Liu, Shuang Li
2024 J jnl
J. Supercomput.
Haiying Liu, Zhihao Li, Kuihua Huang, Rui Wang, Guangquan Cheng, Tiexiang Li
2024 B conf
TrustCom
Haiying Liu, Ruichun Gu, Jingyu Wang, Xiaolin Zhang, Bolin Zhang, Xuebao Li
2024 C conf
CoDIT
Bo Ren, Haiying Liu, Yantong Li
2024 J jnl
Clust. Comput.
Haiying Liu, Kun Wu, Kuihua Huang, Guangquan Cheng, Rui Wang, Guohua Liu
2024 J jnl
IEEE Signal Process. Lett.
Yunhao Chi, Lu Yang, Fanchang Hao, Haiying Liu
2024 J jnl
IEEE Access
Haiying Liu, Haiyan Lin, Xianhui Wang
2023 J jnl
Unmanned Syst.
Haiying Liu, Jie Chen, Jianxin Feng, Haiping Zhao
2023 J jnl
Int. J. Agric. Environ. Inf. Syst.
Haiying Liu, Yongcai Lai, Zhenhua Xu, Zhonliang Yang, Yanmin Yu, Ping Yan
2023 J jnl
Expert Syst. Appl.
Chao Wu, Wenjie Bi, Haiying Liu
2023 J jnl
Remote. Sens.
HongChun Zhu, Zhiwei Lu, Chao Zhang, Yanrui Yang, Guocan Zhu, Yining Zhang, Haiying Liu
2023 J jnl
Sensors
Junmei Guo, Xingchen Liu, Lingyun Bi, Haiying Liu, Haitong Lou
2022 J jnl
Remote. Sens.
Shengjia Zhang, HongChun Zhu, Jie Li, Yanrui Yang, Haiying Liu
2022 J jnl
Trans. Inst. Meas. Control
Yang Zhao, Yan Li, Fangfang Zhang, Haiying Liu
2022 conf
ICISCAE
Lingyun Bi, Lixia Deng, Hongquan Li, Haiying Liu
2022 J jnl
Pattern Recognit.
Yuwen Huang, Gongping Yang, Kuikui Wang, Haiying Liu, Yilong Yin
2022 J jnl
Sensors
Haiying Liu, Fengqian Sun, Jason Gu, Lixia Deng
2022 conf
ICISCAE
Tao Sun, Haonan Chen, Xuehu Duan, Haitong Lou, Haiying Liu
2021 J jnl
ISPRS Int. J. Geo Inf.
YueXue Xu, Shengjia Zhang, Jinyu Li, Haiying Liu, HongChun Zhu
2021 J jnl
IEEE Trans. Inf. Forensics Secur.
Yuwen Huang, Gongping Yang, Kuikui Wang, Haiying Liu, Yilong Yin
2021 J jnl
J. Signal Process. Syst.
Haiying Liu, Ruizhe Ma, Daiyi Li, Li Yan, Zongmin Ma
2021 J jnl
J. Comput. Sci. Technol.
Yuwen Huang, Gongping Yang, Kuikui Wang, Haiying Liu, Yilong Yin
2020 J jnl
Int. J. Mach. Learn. Cybern.
Haiying Liu, Gongping Yang, Yilong Yin
2020 J jnl
Comput. Math. Methods Medicine
Yuansheng Rao, Haiying Liu, Xiaojuan Yan, Jianhong Wang
2020 J jnl
Medical Biol. Eng. Comput.
Sen Zhao, Haiying Liu, Yonghe Li, Yang Song, Wei Wang, Chunqiu Zhang
2020 J jnl
IEEE Access
Jianxin Feng, Rongni Yang, Haiying Liu, Biao Xu
2020 J jnl
Pattern Recognit. Lett.
Jingxiao Xu, Gongping Yang, Kuikui Wang, Yuwen Huang, Haiying Liu, Yilong Yin
2019 J jnl
Int. J. Robotics Autom.
Chaojiong Huang, Umar Farooq, Haiying Liu, Jason Gu, Jun Luo
2019 conf
ROBIO
Huakang Chen, Pengju Chen, Haiying Liu, Jason Gu, Lixia Deng, Juanting Zhou, Huiyuan Zhou
2019 J jnl
Sci. China Inf. Sci.
Haiying Liu, Gongping Yang, Lu Yang, Kun Su, Yilong Yin
2019 J jnl
IEEE CAA J. Autom. Sinica
Chaojiong Huang, Jason Gu, Haiying Liu, Yuansheng Lu, Jun Luo
2019 J jnl
CoRR
Mingming Lu, Kunfang Zhang, Haiying Liu, Naixue Xiong
2019 conf
ROBIO
Juanting Zhou, Pengju Chen, Haiying Liu, Jason Gu, Hui Zhang, Huakang Chen, Huiyuan Zhou
2019 J jnl
Neurocomputing
Haiying Liu, Gongping Yang, Lu Yang, Yilong Yin
2019 conf
ACAI
Haiying Liu, Meng Jian, Bowen Yang, Heng Zhang
2019 J jnl
IEEE Access
Lu Yang, Gongping Yang, Kuikui Wang, Haiying Liu, Xiaoming Xi, Yilong Yin
2018 conf
ADHIP
Fangyi Zhang, Gang Zhao, Haiying Liu, Wang Qin
2018 ch.
Multimedia Tools and Applications for Environmental & Biodiversity Informatics
Haiying Liu, Mike Kobernus, Mirjam Fredriksen, Yaela Golumbic, Johanna Robinson
2018 conf
iThings/GreenCom/CPSCom/SmartData
Junmei Guo, Haiying Liu, Yizhong Luan, Yun Wu
2018 J jnl
Int. J. Spatial Data Infrastructures Res.
Haiying Liu, Arne J. Berre, Mike Kobernus, Mirjam Fredriksen, Richard Rombouts, Andrei Tamlin, Tom Cole-Hunter, Leonardo Santiago, Alena Bartonova
2018 J jnl
IEEE Access
Haiying Liu, Lu Yang, Gongping Yang, Yilong Yin
2018 conf
iThings/GreenCom/CPSCom/SmartData
Haiying Liu, Junmei Guo, Fangzhou Xu, Lixia Deng, Hui Zhang
2018 conf
ADHIP
Xinghua Lu, Gang Zhao, Haiying Liu, Fangyi Zhang
2018 conf
ICIA
Minglei Li, Tao Sun, Haiying Liu
2018 conf
ICBBS
Xige Ma, Haiying Liu
2018 conf
ICIA
Hui Zhang, Haiying Liu, Lixia Deng, Pei Wang, Xuewen Rong, Yibin Li, Bin Li, Haiyan Wang
2018 conf
iThings/GreenCom/CPSCom/SmartData
Huiyuan Zhou, Shubo Qiu, Haiying Liu, Chao Huang
2018 conf
ICIA
Chao Huang, Haiying Liu, Huiyuan Zhou, Valentin Acs
2018 J jnl
Signal Process.
Haiying Liu, Lifang Wu, Dai Zhang, Meng Jian, Xiuzhen Zhang
2018 conf
PRCV (3)
Lifang Wu, Dai Zhang, Meng Jian, Bowen Yang, Haiying Liu
2018 conf
ICMHI
Xige Ma, Haiying Liu, Sen Zhao, Chunqiu Zhang
2017 conf
CCBR
Haiying Liu, Lingfei Song, Gongping Yang, Lu Yang, Yilong Yin
2017 conf
ICIMCS
Lifang Wu, Lei Zhang, Meng Jian, Dai Zhang, Haiying Liu
2017 J jnl
Des. Codes Cryptogr.
Haiying Liu, Keqin Feng, Rongquan Feng
2016 conf
IFIN/ISEM@INSCI
Haiying Liu, Irene Eleta, Mike Kobernus, Tom Cole-Hunter
2016 J jnl
IEEE Access
Haiying Liu, Jason Gu, Max Q.-H. Meng, Wu-Sheng Lu
2016 conf
ICIA
Haiying Liu, Jason Gu, Chaojiong Huang
2016 conf
ICIMCS
Lifang Wu, Dai Zhang, Xiuzhen Zhang, Yuchen Jing, Haiying Liu, Chang Wen Chen
2015 conf
Symposium on Computer Animation
Haiying Liu, Zhixin Yan, Robert W. Lindeman, Gangyi Ding, Tianyu Huang
2015 J jnl
CoRR
Haiying Liu, Keqin Feng, Rongquan Feng
2015 J jnl
Intell. Autom. Soft Comput.
Jiahua Wang, Jingjing Cheng, Haiying Liu, Zhihui Tang, Donghai Han
2014 conf
EnviroInfo
Núria Castell, Haiying Liu, Mike Kobernus, Arne J. Berre, Josef Noll, Erol Cagatay, Reidun Gangdal
2013 J jnl
Int. J. Inf. Acquis.
Haiying Liu, Yu Cheng, Max Q.-H. Meng
2013 conf
ICIA
Min Li, Qiqiang Li, Haiying Liu
2013 conf
ICIA
Binpeng Wang, Haiying Liu
2012 conf
ICAL
Yanan Fu, David W. Zhang, Haiying Liu, Max Q.-H. Meng
2012 B conf
ICMI
Tianyu Huang, Haiying Liu, Gangyi Ding
2011 B conf
ICIP
Juan Song, Keyan Wang, Haiying Liu, Yunsong Li, Chengke Wu
2010 J jnl
Int. J. Pattern Recognit. Artif. Intell.
Xiaochun Cao, Lin Wu, Zeeshan Rasheed, Haiying Liu, Tae Eun Choe, Feng Guo, Niels Haering
2010 conf
IIH-MSP
Haiying Liu, Yunsong Li, Song Xiao, Chengke Wu
2010 J jnl
J. Comput. Chem.
Haiying Liu, Xiaohua Chen, Yuxiang Bu
2009 J jnl
Discret. Math.
Zhiping Wang, Yutang Zou, Haiying Liu, Zhongtuo Wang
2008 conf
ICDSC
Zeeshan Rasheed, Xiaochun Cao, Khurram Shafique, Haiying Liu, Li Yu, Mun Wai Lee, Krishnan Ramnath, Tae Eun Choe, Omar Javed, Niels Haering
2008 B conf
ICPR
Xiaochun Cao, Zeeshan Rasheed, Haiying Liu, Niels Haering
2007 conf
ICASSP (1)
Haiying Liu, Rama Chellappa
2007 J jnl
Int. J. Wavelets Multiresolution Inf. Process.
Haiying Liu, Prabir Bhattacharya
2004 conf
CARS
Walter A. Hall, Haiying Liu, Charles L. Truwit
2003 J jnl
IEEE Trans. Image Process.
Haiying Liu, Rama Chellappa, Azriel Rosenfeld
2003 J jnl
IEEE Trans. Medical Imaging
Thomas Hartkens, Derek L. G. Hill, Andy D. Castellano-Smith, David J. Hawkes, Calvin R. Maurer Jr., Alastair J. Martin, Walter A. Hall, Haiying Liu, Charles L. Truwit
2002 conf
ICPR (1)
Haiying Liu, Rama Chellappa, Azriel Rosenfeld
2002 conf
Image-Guided Procedures
Haiying Liu, Walter A. Hall, Charles L. Truwit
2002 conf
Image-Guided Procedures
Haiying Liu, Sean O. Casey, Edward Michel, Charles L. Truwit
2002 B conf
Image Processing
Andrew D. Castellano-Smith, Thomas Hartkens, Julia Anne Schnabel, D. Rodney Hose, Haiying Liu, Walter A. Hall, Charles L. Truwit, David John Hawkes, Derek L. G. Hill
2002 Misc conf
ICASSP
Haiying Liu, Azriel Rosenfeld, Rama Chellappa
2002 conf
MICCAI (2)
Thomas Hartkens, Derek L. G. Hill, Andy D. Castellano-Smith, David J. Hawkes, Calvin R. Maurer Jr., Alastair J. Martin, Walter A. Hall, Haiying Liu, Charles L. Truwit
2002 conf
Image-Guided Procedures
Haiying Liu, Weu-Tek Tsai, Gino Canessa, John C. Canessa
2001 A conf
MICCAI
Julia A. Schnabel, Daniel Rueckert, Marcel Quist, Jane M. Blackall, Andy D. Castellano-Smith, Thomas Hartkens, Graeme P. Penney, Walter A. Hall, Haiying Liu, Charles L. Truwit, Frans A. Gerritsen, Derek L. G. Hill, David J. Hawkes
2001 A conf
MICCAI
Andy D. Castellano-Smith, Thomas Hartkens, Julia A. Schnabel, D. Rodney Hose, Haiying Liu, Walter A. Hall, Charles L. Truwit, David J. Hawkes, Derek L. G. Hill
2001 conf
Image-Guided Procedures
Haiying Liu, Robert E. Maxwell, Charles L. Truwit
2001 conf
Image-Guided Procedures
Haiying Liu, Chialei Chin
2001 conf
Image-Guided Procedures
Haiying Liu, Walter A. Hall, Charles L. Truwit
2000 J jnl
Pattern Recognit. Lett.
Prabir Bhattacharya, Haiying Liu, Azriel Rosenfeld, W. Scott Thompson
2000 B conf
ICPR
Haiying Liu, Prabir Bhattacharya
1998 J jnl
IEEE Trans. Medical Imaging
Calvin R. Maurer Jr., Derek L. G. Hill, Alastair J. Martin, Haiying Liu, M. McCue, Daniel Rueckert, David Lloret, Walter A. Hall, Robert E. Maxwell, David J. Hawkes, Charles L. Truwit
1998 J jnl
IEEE Trans. Medical Imaging
Haiying Liu, Charles L. Truwit
tests/integration/test_js_extractors.py
← Index tests/integration/test_js_extractors.py python
"""
Integration tests for JavaScript-specific extractors.
"""
import sys
import types
from unittest.mock import MagicMock

# Mock the oscrypto/signify/certvalidator chain before any redb imports.
# oscrypto fails on environments with OpenSSL 3.x due to version detection.
# JS extractors don't use PE signatures so this is safe.
_MOCK_MODULES = [
    # oscrypto/signify chain (PE signature analysis)
    "oscrypto", "oscrypto.errors", "oscrypto._openssl",
    "oscrypto._openssl._libcrypto", "oscrypto._openssl._libcrypto_ctypes",
    "oscrypto._openssl.util", "oscrypto.util", "oscrypto.kdf",
    "oscrypto._asymmetric", "oscrypto.asymmetric",
    "certvalidator", "certvalidator.validate",
    "signify", "signify.fingerprinter", "signify.authenticode",
    "signify.authenticode.signed_pe", "signify.authenticode.authroot",
    "signify.pkcs7", "signify.pkcs7.signeddata", "signify.pkcs7.signerinfo",
    "signify.x509", "signify.x509.certificates", "signify.x509.context",
    # hashing libs used by HashExtractor
    "ppdeep", "tlsh",
    # binary format libs (not needed for JS)
    "dotnetfile", "lief",
    "flare_floss", "yara_x",
    "machofile",
]
for _mod in _MOCK_MODULES:
    if _mod not in sys.modules:
        try:
            __import__(_mod)
        except (ImportError, Exception):
            sys.modules[_mod] = MagicMock()

import pytest
from pathlib import Path
from unittest.mock import patch

pytestmark = [pytest.mark.integration, pytest.mark.js]

TEST_FILES_DIR = Path(__file__).parent.parent.parent / "test_files"

JS_MALICIOUS_PATH = str(TEST_FILES_DIR / "test_malicious.js")


@pytest.fixture
def js_malicious_path():
    path = TEST_FILES_DIR / "test_malicious.js"
    if not path.exists():
        pytest.skip(f"Test file not found: {path}")
    return str(path)


# ============================================================================
# JSFeaturesExtractor Tests
# ============================================================================

class TestJSFeaturesExtractor:

    def test_extract_valid_js(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor
        from redb.models.dataclasses import JSFeatures

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            assert result is not None
            assert isinstance(result, JSFeatures)

    def test_entropy_is_reasonable(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            assert result.text_entropy > 0.0
            assert result.text_entropy < 8.0

    def test_detects_obfuscation(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            # The fixture is dense `\xHH` hex escapes (>5% of source) plus
            # text entropy >5.0 — two strong signals, which clears the
            # strong-signal gate the heuristic now requires.
            assert result.is_likely_obfuscated is True
            assert result.obfuscation_score >= 40
            assert len(result.obfuscation_techniques) > 0
            # `obfuscator_name` is None when js-x-ray isn't installed in the
            # test environment, and a recognised family name when it is.
            assert result.obfuscator_name is None or isinstance(
                result.obfuscator_name, str
            )

    def test_detects_eval_usage(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            assert result.eval_count >= 1
            assert result.fromcharcode_count >= 1
            assert result.atob_count >= 1

    def test_detects_wscript_environment(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            assert result.detected_environment == "wscript"
            assert result.script_type == "wscript"

    def test_line_metrics(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            assert result.line_count > 0
            assert result.char_count > 0
            assert result.max_line_length > 0
            assert result.avg_line_length > 0

    def test_export_data_format(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(js_malicious_path, mock_logger)
            extractor.extract()
            export = extractor.prepare_export_data('ClickHouseExporter')

            assert export is not None
            data, col_names, col_types = export
            assert len(col_names) == len(col_types)
            assert len(data) == 1
            assert len(data[0]) == len(col_names)

    def test_clickhouse_table_name(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(js_malicious_path, mock_logger)
            assert extractor.get_clickhouse_table() == "redb_js_features"

    def test_empty_file_returns_none(self, tmp_path, mock_logger):
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        empty_js = tmp_path / "empty.js"
        empty_js.write_text("")

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(str(empty_js), mock_logger)
            result = extractor.extract()
            assert result is None

    def test_pattern_counts_are_case_insensitive(self, tmp_path, mock_logger):
        """eval_count / atob_count come from js_patterns.PATTERNS, which are
        compiled with re.IGNORECASE. Capitalised forms must still be counted."""
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        mixed_case_js = tmp_path / "mixed.js"
        mixed_case_js.write_text(
            'EVAL("1+1");\n'
            'eval("2+2");\n'
            'ATOB("YWJj");\n'
        )

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(str(mixed_case_js), mock_logger)
            result = extractor.extract()
            assert result is not None
            # Both `EVAL(` and `eval(` should be counted.
            assert result.eval_count == 2
            assert result.atob_count == 1

    # ----- script_type / detected_environment classification ----------------
    # First-match-wins; ordering is what makes the HTA-vs-WScript split work,
    # so we pin one positive case per value and one orthogonality case.

    @pytest.mark.parametrize("source,expected_type", [
        ("#@~^abcdef==^#~@", "jse"),
        ("<?xml version='1.0'?>\n<job><script language='JScript'>x=1;</script></job>", "wsf"),
        ("<html><head><hta:application id='x'/></head><script>x=1</script></html>", "hta"),
        ("<html><body><script>alert(1)</script></body></html>", "embedded_html"),
        ("var s = new ActiveXObject('WScript.Shell');", "wscript"),
        ("import {foo} from 'bar';\nexport const x = 1;", "esm"),
        ("import 'side-effect.js';\nconst x = 1;", "esm"),
        ("const fs = require('fs');\nmodule.exports = {};", "node_module"),
        ("function add(a, b) { return a + b; }", "standalone"),
    ])
    def test_script_type_classification(self, tmp_path, mock_logger, source, expected_type):
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        js_file = tmp_path / "sample.js"
        js_file.write_text(source)
        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(str(js_file), mock_logger, source=source)
            result = extractor.extract()
            assert result is not None
            assert result.script_type == expected_type

    @pytest.mark.parametrize("source,expected_env", [
        ("var s = new ActiveXObject('WScript.Shell');", "wscript"),
        ("chrome.runtime.onMessage.addListener(()=>{}); chrome.tabs.query({});", "browser_extension"),
        ("self.addEventListener('fetch', e => e.respondWith(caches.match(e.request)));", "service_worker"),
        ("const text = await Deno.readTextFile('a.txt');", "deno"),
        ("const fs = require('fs'); process.env.X;", "node"),
        ("document.getElementById('x'); window.location;", "browser"),
        ("function add(a, b) { return a + b; }", "unknown"),
    ])
    def test_environment_classification(self, tmp_path, mock_logger, source, expected_env):
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        js_file = tmp_path / "sample.js"
        js_file.write_text(source)
        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(str(js_file), mock_logger, source=source)
            result = extractor.extract()
            assert result is not None
            assert result.detected_environment == expected_env

    def test_hta_with_wscript_keeps_format_and_runtime_orthogonal(self, tmp_path, mock_logger):
        """HTA droppers commonly call WScript APIs. script_type captures the
        file format (`hta`) and detected_environment captures the runtime
        surface (`wscript`); both axes must survive a single sample."""
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        source = (
            "<html><head><hta:application id='x'/></head>"
            "<script>new ActiveXObject('WScript.Shell').Run('cmd.exe');</script></html>"
        )
        js_file = tmp_path / "sample.hta"
        js_file.write_text(source)
        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(str(js_file), mock_logger, source=source)
            result = extractor.extract()
            assert result.script_type == "hta"
            assert result.detected_environment == "wscript"

    def test_bare_exports_property_is_not_node_module(self, tmp_path, mock_logger):
        """Regression: `module.exports` and `require(` mark a CommonJS module,
        but a bare `exports.` substring (e.g. `obj.exports.foo`) should not."""
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        source = "var x = obj.exports.foo;\nfunction f() { return 1; }\n"
        js_file = tmp_path / "sample.js"
        js_file.write_text(source)
        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSFeaturesExtractor(str(js_file), mock_logger, source=source)
            result = extractor.extract()
            assert result.script_type == "standalone"


# ============================================================================
# JSSuspiciousAPIsExtractor Tests
# ============================================================================

class TestJSSuspiciousAPIsExtractor:

    def test_extract_finds_apis(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_suspicious_apis import JSSuspiciousAPIsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSSuspiciousAPIsExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            assert result is not None
            assert len(result) > 0

    def test_detects_eval(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_suspicious_apis import JSSuspiciousAPIsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSSuspiciousAPIsExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            api_names = [f['api_name'] for f in result]
            assert "eval" in api_names

    def test_detects_wscript_apis(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_suspicious_apis import JSSuspiciousAPIsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSSuspiciousAPIsExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            api_names = [f['api_name'] for f in result]
            assert "WScript.CreateObject" in api_names
            assert "WScript.Shell" in api_names

    def test_detects_network_apis(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_suspicious_apis import JSSuspiciousAPIsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSSuspiciousAPIsExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            api_names = [f['api_name'] for f in result]
            assert "XMLHttpRequest" in api_names
            assert "fetch" in api_names

    def test_detects_registry_access(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_suspicious_apis import JSSuspiciousAPIsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSSuspiciousAPIsExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            api_names = [f['api_name'] for f in result]
            assert "RegWrite" in api_names

    def test_categories_are_valid(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_suspicious_apis import JSSuspiciousAPIsExtractor, SUSPICIOUS_APIS

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSSuspiciousAPIsExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            valid_categories = set(SUSPICIOUS_APIS.keys())
            for finding in result:
                assert finding['api_category'] in valid_categories
                assert finding['call_count'] > 0
                assert len(finding['line_numbers']) > 0

    def test_export_data_format(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_suspicious_apis import JSSuspiciousAPIsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSSuspiciousAPIsExtractor(js_malicious_path, mock_logger)
            extractor.extract()
            export = extractor.prepare_export_data('ClickHouseExporter')

            assert export is not None
            data, col_names, col_types = export
            assert len(col_names) == len(col_types)
            assert len(data) > 0
            assert len(data[0]) == len(col_names)

    def test_clickhouse_table_name(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_suspicious_apis import JSSuspiciousAPIsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSSuspiciousAPIsExtractor(js_malicious_path, mock_logger)
            assert extractor.get_clickhouse_table() == "redb_js_suspicious_apis"

    def test_benign_js_has_no_findings(self, tmp_path, mock_logger):
        from redb.extractors.js_extractors.js_suspicious_apis import JSSuspiciousAPIsExtractor

        benign_js = tmp_path / "benign.js"
        benign_js.write_text("function add(a, b) { return a + b; }\nconsole.log(add(1, 2));")

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSSuspiciousAPIsExtractor(str(benign_js), mock_logger)
            result = extractor.extract()
            assert result is None

    def test_pattern_match_is_case_insensitive(self, tmp_path, mock_logger):
        """Patterns in js_patterns.PATTERNS are compiled with re.IGNORECASE.
        A capitalised `EVAL(` is not actually valid JS at runtime but obfuscated
        droppers occasionally use stringly-built names; verifying the match
        survives capitalisation guards against an accidental flag regression.
        """
        from redb.extractors.js_extractors.js_suspicious_apis import JSSuspiciousAPIsExtractor

        upper_js = tmp_path / "upper.js"
        upper_js.write_text('var x = EVAL("1+1");\nFETCH("http://evil.test");')

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSSuspiciousAPIsExtractor(str(upper_js), mock_logger)
            result = extractor.extract()
            assert result is not None
            api_names = {f['api_name'] for f in result}
            assert "eval" in api_names
            assert "fetch" in api_names


# ============================================================================
# JSStringsExtractor Tests
# ============================================================================

class TestJSStringsExtractor:

    def test_extract_finds_strings(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_strings import JSStringsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSStringsExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            assert result is not None
            assert len(result) > 0

    def test_decodes_hex_strings(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_strings import JSStringsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSStringsExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            hex_decoded = [s for s in result if s['string_encoding'] == 'hex']
            assert len(hex_decoded) > 0

    def test_decodes_charcode(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_strings import JSStringsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSStringsExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            charcode_decoded = [s for s in result if s['string_encoding'] == 'charcode']
            assert len(charcode_decoded) > 0
            assert any('powershell' in s['string'] for s in charcode_decoded)

    def test_decodes_base64(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_strings import JSStringsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSStringsExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            b64_decoded = [s for s in result if s['string_encoding'] == 'base64']
            assert len(b64_decoded) > 0

    def test_reconstructs_concatenated_strings(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_strings import JSStringsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSStringsExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            concat = [s for s in result if s['string_encoding'] == 'concat']
            assert len(concat) > 0

    def test_strings_have_standard_fields(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_strings import JSStringsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSStringsExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            for s in result:
                assert s['string_entropy'] >= 0
                assert s['string_length'] > 0
                assert s['string_raw_length'] > 0
                assert s['string_offset'] > 0
                assert len(s['string']) > 0
                assert len(s['string_raw']) > 0

    def test_export_data_format(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_strings import JSStringsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSStringsExtractor(js_malicious_path, mock_logger)
            extractor.extract()
            export = extractor.prepare_export_data('ClickHouseExporter')

            assert export is not None
            data, col_names, col_types = export
            assert len(col_names) == len(col_types)
            assert len(data) > 0
            assert len(data[0]) == len(col_names)

    def test_writes_to_shared_strings_table(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_strings import JSStringsExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSStringsExtractor(str(js_malicious_path), mock_logger)
            assert extractor.get_clickhouse_table() == "code_binja_strings_raw"

    def test_line_numbers_correct_in_multiline_source(self, tmp_path, mock_logger):
        """Pin the bisect-based _find_line_number against a multi-line source.

        Builds a JS file where each encoded literal sits on a known line, then
        verifies the `string_offset` (1-indexed line number) the extractor
        emits matches the planted line. Regression guard for the swap from the
        historical O(N) `source[:start].count("\\n")` to O(log L) bisect.
        """
        from redb.extractors.js_extractors.js_strings import JSStringsExtractor

        # 800 filler lines, then a known hex-escape literal on line 801
        # (1-indexed), then 100 more filler, then a charcode call on line 902.
        prefix = "\n".join([f"function f{i}() {{ return {i}; }}" for i in range(800)])
        line_801 = 'var h = "\\x68\\x74\\x74\\x70\\x73\\x3a\\x2f\\x2f"'
        middle = "\n".join([f"var v{i} = {i};" for i in range(100)])
        line_902 = 'var c = String.fromCharCode(112, 111, 119, 101, 114)'
        source = "\n".join([prefix, line_801, middle, line_902]) + "\n"

        js_path = tmp_path / "multiline.js"
        js_path.write_text(source)

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSStringsExtractor(str(js_path), mock_logger)
            result = extractor.extract()

            assert result is not None
            by_encoding = {}
            for f in result:
                by_encoding.setdefault(f['string_encoding'], []).append(f)

            assert any(f['string_offset'] == 801 for f in by_encoding.get('hex', [])), (
                f"hex literal should be reported on line 801; got "
                f"{[f['string_offset'] for f in by_encoding.get('hex', [])]}"
            )
            assert any(f['string_offset'] == 902 for f in by_encoding.get('charcode', [])), (
                f"charcode call should be reported on line 902; got "
                f"{[f['string_offset'] for f in by_encoding.get('charcode', [])]}"
            )

    def test_publishes_findings_to_context(self, js_malicious_path, mock_logger):
        """JSStringsExtractor must mirror its findings onto JSContext.decoded_strings
        so post-loop consumers (the IOC plumbing in workers.py) can read the
        decoded strings without holding a reference to this extractor."""
        from redb.extractors.js_extractors.js_strings import JSStringsExtractor
        from redb.extractors.js_extractors.js_context import JSContext

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            ctx = JSContext.from_path(js_malicious_path, log=mock_logger)
            assert ctx.decoded_strings is None  # not populated until extract runs

            extractor = JSStringsExtractor(js_malicious_path, mock_logger, context=ctx)
            findings = extractor.extract()

            assert ctx.decoded_strings is findings, (
                "ctx.decoded_strings should reference the same list returned "
                "by extract(), not a copy"
            )


# ============================================================================
# JSDeobfuscationExtractor Tests
# ============================================================================

class TestJSDeobfuscationExtractor:

    def test_extract_with_jsbeautifier_fallback(self, js_malicious_path, mock_logger):
        """Should fall back to jsbeautifier when webcrack is not installed."""
        from redb.extractors.js_extractors.js_deobfuscation import JSDeobfuscationExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSDeobfuscationExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            if result is not None:
                assert result['deobfuscator_used'] is not None
                assert result['deobfuscation_successful'] is True
                assert result['original_size'] > 0
                assert result['deobfuscated_size'] > 0
                assert len(result['deobfuscated_sha256']) == 64

    def test_export_data_format(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_deobfuscation import JSDeobfuscationExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSDeobfuscationExtractor(js_malicious_path, mock_logger)
            result = extractor.extract()

            if result is not None:
                export = extractor.prepare_export_data('ClickHouseExporter')
                assert export is not None
                data, col_names, col_types = export
                assert len(col_names) == len(col_types)
                assert len(data[0]) == len(col_names)

    def test_clickhouse_table_name(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_deobfuscation import JSDeobfuscationExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSDeobfuscationExtractor(js_malicious_path, mock_logger)
            assert extractor.get_clickhouse_table() == "redb_js_deobfuscation"

    def test_runs_deobfuscator_without_separate_probe(self, js_malicious_path, mock_logger):
        """The deobfuscator binary is invoked directly — no separate
        `--help` probe per sample. When the binary is missing, the
        FileNotFoundError is swallowed silently (no error log) and the
        extractor falls through to jsbeautifier.
        """
        from redb.extractors.js_extractors import js_deobfuscator as deob_helper
        from redb.extractors.js_extractors.js_deobfuscation import (
            JSDeobfuscationExtractor,
        )

        # _check_tool_available must be gone entirely.
        assert not hasattr(JSDeobfuscationExtractor, "_check_tool_available")

        popen_calls = []

        def fake_popen(*args, **kwargs):
            popen_calls.append(args[0] if args else kwargs.get("args"))
            raise FileNotFoundError(2, "No such file or directory", "fake-deobfuscator")

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'), \
                patch.object(deob_helper.subprocess, 'Popen', fake_popen):
            extractor = JSDeobfuscationExtractor(js_malicious_path, mock_logger)
            extractor.extract()

        # Exactly one Popen attempt — the actual run. No extra probe call.
        assert len(popen_calls) == 1, (
            f"expected one Popen call (the deobfuscator run), got "
            f"{len(popen_calls)}: {popen_calls}"
        )
        # FileNotFoundError must not have been logged at error level.
        for call in mock_logger.error.call_args_list:
            joined = " ".join(str(a) for a in call.args)
            assert "not found" not in joined.lower(), (
                f"missing-binary case must not log at error level; got {call}"
            )


# ============================================================================
# JSContentExtractor Tests
# ============================================================================

class TestJSContentExtractor:
    """JSContentExtractor persists raw + normalised text into code_text_content."""

    def test_extract_populates_both_columns(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_content import JSContentExtractor
        from redb.extractors.js_extractors.js_context import JSContext

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            ctx = JSContext.from_path(js_malicious_path, log=mock_logger)
            extractor = JSContentExtractor(js_malicious_path, mock_logger, context=ctx)
            result = extractor.extract()

            assert result is not None
            assert result['content_type'] == 'javascript'
            assert result['text_raw']
            # jsbeautifier should normalise even non-obfuscated input, so
            # against the bundled test sample we expect a normalised text.
            assert result['text_normalized'] is not None
            assert result['normalizer_used'] is not None

    def test_normalized_is_null_when_deobfuscator_returns_nothing(
        self, js_malicious_path, mock_logger
    ):
        """When neither webcrack nor jsbeautifier produces output the
        normalized columns must be NULL — distinguishes failure from a
        legitimate empty result."""
        from redb.extractors.js_extractors.js_content import JSContentExtractor
        from redb.extractors.js_extractors.js_context import JSContext
        from redb.extractors.js_extractors import js_deobfuscator as deob_helper

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'), \
                patch.object(deob_helper, 'deobfuscate', lambda src, log: (None, None)):
            ctx = JSContext.from_path(js_malicious_path, log=mock_logger)
            extractor = JSContentExtractor(js_malicious_path, mock_logger, context=ctx)
            result = extractor.extract()

            assert result is not None
            assert result['text_raw']
            assert result['text_normalized'] is None
            assert result['normalizer_used'] is None

    def test_does_not_run_its_own_deobfuscation(self, js_malicious_path, mock_logger):
        """JSContentExtractor reads from JSContext.deobfuscated, which is a
        cached_property. Two extractors sharing one context must trigger
        deobfuscation at most once across both extracts."""
        from redb.extractors.js_extractors.js_content import JSContentExtractor
        from redb.extractors.js_extractors.js_deobfuscation import (
            JSDeobfuscationExtractor,
        )
        from redb.extractors.js_extractors.js_context import JSContext
        from redb.extractors.js_extractors import js_deobfuscator as deob_helper

        calls = {"n": 0}

        def counting_deobfuscate(source, log):
            calls["n"] += 1
            return None, None

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'), \
                patch.object(deob_helper, 'deobfuscate', counting_deobfuscate):
            ctx = JSContext.from_path(js_malicious_path, log=mock_logger)
            JSDeobfuscationExtractor(js_malicious_path, mock_logger, context=ctx).extract()
            JSContentExtractor(js_malicious_path, mock_logger, context=ctx).extract()

            assert calls["n"] == 1, (
                f"deobfuscate() must be called once across both extractors "
                f"sharing a context, got {calls['n']}"
            )

    def test_export_data_format(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_content import JSContentExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSContentExtractor(js_malicious_path, mock_logger)
            extractor.extract()
            data, col_names, col_types = extractor.prepare_export_data(
                'ClickHouseExporter'
            )
            assert col_names == [
                'sha256', 'content_type', 'text_raw', 'text_normalized',
                'normalizer_used', 'analysis_date',
            ]
            assert len(col_names) == len(col_types) == len(data[0])
            assert data[0][1] == 'javascript'

    def test_clickhouse_table_name(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_content import JSContentExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            extractor = JSContentExtractor(js_malicious_path, mock_logger)
            assert extractor.get_clickhouse_table() == 'code_text_content'


# ============================================================================
# Source Sharing Tests
# ============================================================================

class TestSourceSharing:
    """Test that source text can be shared across extractors."""

    def test_shared_source_produces_same_results(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            with open(js_malicious_path, 'r') as f:
                source = f.read()

            ext_shared = JSFeaturesExtractor(js_malicious_path, mock_logger, source=source)
            ext_fresh = JSFeaturesExtractor(js_malicious_path, mock_logger)

            r_shared = ext_shared.extract()
            r_fresh = ext_fresh.extract()

            assert r_shared.text_entropy == r_fresh.text_entropy
            assert r_shared.eval_count == r_fresh.eval_count
            assert r_shared.obfuscation_score == r_fresh.obfuscation_score


# ============================================================================
# JSContext Sharing Tests
# ============================================================================

class TestJSContextSharing:
    """Verify the shared JSContext threads through every extractor cleanly:
    same outputs as independent runs, expensive work computed exactly once."""

    def test_shared_context_produces_same_results(self, js_malicious_path, mock_logger):
        from redb.extractors.js_extractors.js_context import JSContext
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor
        from redb.extractors.js_extractors.js_suspicious_apis import (
            JSSuspiciousAPIsExtractor,
        )

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            ctx = JSContext.from_path(js_malicious_path, log=mock_logger)

            shared_features = JSFeaturesExtractor(
                js_malicious_path, mock_logger, context=ctx
            ).extract()
            shared_apis = JSSuspiciousAPIsExtractor(
                js_malicious_path, mock_logger, context=ctx
            ).extract()

            fresh_features = JSFeaturesExtractor(
                js_malicious_path, mock_logger
            ).extract()
            fresh_apis = JSSuspiciousAPIsExtractor(
                js_malicious_path, mock_logger
            ).extract()

            # Features dataclass must compare equal across paths.
            assert shared_features == fresh_features
            # Suspicious-API findings: same set of api_name/category pairs.
            shared_pairs = {(f['api_name'], f['api_category']) for f in shared_apis}
            fresh_pairs = {(f['api_name'], f['api_category']) for f in fresh_apis}
            assert shared_pairs == fresh_pairs

    def test_context_scan_is_computed_once(self, js_malicious_path, mock_logger):
        """Two extractors sharing one JSContext must each trigger their backing
        scan at most once.

        The context exposes two cached scan results: `scan` (raw source) and
        `scan_deobfuscated` (deobfuscated text, populated only by the dual-pass
        consumers in JSSuspiciousAPIsExtractor / JSDeobfuscationExtractor).
        Both are cached_property — repeating the access from a second extractor
        must not re-run scan_source. Whether the deobfuscated scan happens at
        all depends on the deobfuscator producing a different text for this
        sample; we accept either outcome and only fail if the per-cache
        single-run invariant is broken.
        """
        from redb.extractors.js_extractors import js_context as js_ctx_module
        from redb.extractors.js_extractors.js_context import JSContext
        from redb.extractors.js_extractors.js_features import JSFeaturesExtractor
        from redb.extractors.js_extractors.js_suspicious_apis import (
            JSSuspiciousAPIsExtractor,
        )

        real_scan = js_ctx_module.scan_source
        call_count = {"n": 0}

        def counting_scan(*args, **kwargs):
            call_count["n"] += 1
            return real_scan(*args, **kwargs)

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'), \
                patch.object(js_ctx_module, 'scan_source', counting_scan):
            ctx = JSContext.from_path(js_malicious_path, log=mock_logger)

            JSFeaturesExtractor(
                js_malicious_path, mock_logger, context=ctx
            ).extract()
            JSSuspiciousAPIsExtractor(
                js_malicious_path, mock_logger, context=ctx
            ).extract()
            after_first_pair = call_count["n"]

            # The raw scan and (optionally) the deobfuscated scan are the only
            # two scan_source calls allowed for this sample.
            assert 1 <= after_first_pair <= 2, (
                f"scan_source must run at most twice (raw + deobf) when context "
                f"is shared, got {after_first_pair}"
            )

            # Re-running the same extractors must not re-trigger any scan: both
            # cached_property results are pinned by the first access.
            JSFeaturesExtractor(
                js_malicious_path, mock_logger, context=ctx
            ).extract()
            JSSuspiciousAPIsExtractor(
                js_malicious_path, mock_logger, context=ctx
            ).extract()

            assert call_count["n"] == after_first_pair, (
                f"cached scans re-ran across extractor instances: "
                f"{after_first_pair} -> {call_count['n']}"
            )

    def test_context_deobfuscation_runs_once(self, js_malicious_path, mock_logger):
        """Multiple consumers reading JSContext.deobfuscated must trigger the
        external deobfuscator at most once. cached_property guarantees this;
        the test pins the contract so a future regression fails loudly."""
        from redb.extractors.js_extractors import js_context as js_ctx_module
        from redb.extractors.js_extractors.js_context import JSContext

        deob_calls = {"n": 0}
        real_helper = None

        def counting_deobfuscate(source, log):
            deob_calls["n"] += 1
            return None, None  # cheap stub: no need to actually run webcrack

        # Patch the import that JSContext.deobfuscated does lazily inside the
        # cached_property. We monkeypatch the helper module directly.
        from redb.extractors.js_extractors import js_deobfuscator as deob_helper

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'), \
                patch.object(deob_helper, 'deobfuscate', counting_deobfuscate):
            ctx = JSContext.from_path(js_malicious_path, log=mock_logger)
            _ = ctx.deobfuscated
            _ = ctx.deobfuscated  # second access must hit the cache
            _ = ctx.deobfuscated  # third access too

            assert deob_calls["n"] == 1, (
                f"deobfuscate() must be called once per context, got "
                f"{deob_calls['n']}"
            )

    def test_context_ast_parsed_once(self, js_malicious_path, mock_logger):
        """Repeated access to JSContext.ast must invoke pyjsparser.parse at
        most once. cached_property guarantees this; the test pins the
        contract so a future regression (e.g. dropping cached_property)
        fails loudly."""
        from redb.extractors.js_extractors.js_context import JSContext

        try:
            import pyjsparser  # noqa: F401
        except ImportError:
            pytest.skip("pyjsparser not installed")

        parse_calls = {"n": 0}
        real_parse = pyjsparser.parse

        def counting_parse(source):
            parse_calls["n"] += 1
            return real_parse(source)

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'), \
                patch.object(pyjsparser, 'parse', counting_parse):
            ctx = JSContext.from_path(js_malicious_path, log=mock_logger)
            _ = ctx.ast
            _ = ctx.ast  # second access: must hit the cache
            _ = ctx.ast  # third access: still cached

            # parse may legitimately be 0 if the source can't be parsed and
            # the exception was swallowed; otherwise it must be exactly 1.
            assert parse_calls["n"] <= 1, (
                f"pyjsparser.parse called {parse_calls['n']} times for one "
                "context — cached_property must hold the parsed AST"
            )


# ============================================================================
# JS IOC Surface Tests (end-to-end)
# ============================================================================

class TestJSIOCSurfaces:
    """workers.py-style assembly of analysis_results: text_raw, text_normalized,
    and decoded strings. End-to-end through IOCExtractorFromResults."""

    def _build_analysis_results(self, ctx, raw_source, sha256):
        """Mirror the dict workers.py builds in the JS IOC block. Kept in a
        helper so the test exercises the same shape the production code uses
        without dragging the full workers.py dispatch into the test harness."""
        import hashlib

        text_raw_entries = [{"content": raw_source, "content_hash": sha256}]
        text_normalized_entries = []
        deobf_text, _ = ctx.deobfuscated
        if deobf_text and deobf_text != raw_source:
            text_normalized_entries.append({
                "content": deobf_text,
                "content_hash": hashlib.sha256(
                    deobf_text.encode("utf-8")
                ).hexdigest(),
            })
        decoded_strings = ctx.decoded_strings or []
        strings_entries = [
            {
                "string": s.get("string", ""),
                "string_offset": s.get("string_offset", 0),
            }
            for s in decoded_strings
        ]
        return {
            "strings": strings_entries,
            "text_raw": text_raw_entries,
            "text_normalized": text_normalized_entries,
        }

    def test_decoded_string_url_lands_via_strings_surface(self, tmp_path, mock_logger):
        """A URL hidden inside a String.fromCharCode(...) call is decoded by
        JSStringsExtractor into a plaintext string; running IOC extraction
        over the decoded `strings` surface picks it up with source_type=string.
        Without the broadened plumbing the URL would be invisible to IOC
        extraction (regex doesn't see it through the charcode encoding in
        the raw source)."""
        from redb.extractors.ioc_extractor.ioc_extractor import IOCExtractorFromResults
        from redb.extractors.ioc_extractor.standalone_ioc_extractor import (
            IOCType, SourceType,
        )
        from redb.extractors.js_extractors.js_context import JSContext
        from redb.extractors.js_extractors.js_strings import JSStringsExtractor

        # String.fromCharCode of "http://hidden.example.com" (25 chars, codes
        # built inline below).
        url = "http://hidden.example.com"
        codes = ", ".join(str(ord(c)) for c in url)
        source = (
            'var benign = "harmless string here";\n'
            f'var u = String.fromCharCode({codes});\n'
        )
        sample_path = tmp_path / "charcode.js"
        sample_path.write_text(source)

        sha256 = "b" * 64

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            ctx = JSContext.from_path(str(sample_path), log=mock_logger)
            JSStringsExtractor(str(sample_path), mock_logger, context=ctx).extract()

            analysis_results = self._build_analysis_results(ctx, ctx.source, sha256)
            iocs = IOCExtractorFromResults(
                analysis_results=analysis_results,
                sha256=sha256,
                log=mock_logger,
            ).extract()

            url_iocs = [i for i in iocs if i.ioc_type == IOCType.URL]
            assert any(
                "hidden.example.com" in i.ioc_value
                and i.source_type == SourceType.STRING
                for i in url_iocs
            ), (
                f"charcode-decoded URL must be picked up via the strings "
                f"surface; got {[(i.ioc_value, i.source_type) for i in url_iocs]}"
            )

    def test_raw_url_lands_via_text_raw(self, tmp_path, mock_logger):
        """A plain URL in the raw source lands with source_type=text_raw, the
        new universal surface (was source_type=decompiled_function before #C)."""
        from redb.extractors.ioc_extractor.ioc_extractor import IOCExtractorFromResults
        from redb.extractors.ioc_extractor.standalone_ioc_extractor import (
            IOCType, SourceType,
        )
        from redb.extractors.js_extractors.js_context import JSContext

        source = 'fetch("https://plain.example.com/c2");\n'
        sample_path = tmp_path / "plain.js"
        sample_path.write_text(source)
        sha256 = "c" * 64

        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'):
            ctx = JSContext.from_path(str(sample_path), log=mock_logger)
            analysis_results = self._build_analysis_results(ctx, ctx.source, sha256)
            iocs = IOCExtractorFromResults(
                analysis_results=analysis_results,
                sha256=sha256,
                log=mock_logger,
            ).extract()

            url_iocs = [i for i in iocs if i.ioc_type == IOCType.URL]
            assert any(
                "plain.example.com" in i.ioc_value
                and i.source_type == SourceType.TEXT_RAW
                for i in url_iocs
            )

    def test_no_text_normalized_entry_when_deobf_equals_raw(self, tmp_path, mock_logger):
        """If the deobfuscator returns text byte-identical to the raw source
        (or returns nothing), the analysis_results dict must not add a
        text_normalized entry — avoids double-scraping the same content."""
        from redb.extractors.js_extractors.js_context import JSContext
        from redb.extractors.js_extractors import js_deobfuscator as deob_helper

        source = 'fetch("https://noop.example.com/");\n'
        sample_path = tmp_path / "noop.js"
        sample_path.write_text(source)
        sha256 = "d" * 64

        # Force the deobfuscator to return identical text — exercises the
        # equality check in workers.py / _build_analysis_results.
        with patch('redb.settings.ELASTIC_BINARIES_COLLECTION', 'test'), \
                patch.object(deob_helper, 'deobfuscate',
                             lambda src, log: (src, "stub")):
            ctx = JSContext.from_path(str(sample_path), log=mock_logger)
            analysis_results = self._build_analysis_results(ctx, ctx.source, sha256)

            assert analysis_results["text_normalized"] == [], (
                "text_normalized must be empty when deobfuscation produced "
                "byte-identical output"
            )