Manish Shrivastava

231 papers A* 6A 1B 15C 6Misc 4Journal 74Unranked 125
YearRankTypeTitle / Venue / Authors
2026 J jnl
CoRR
Saumitra Yadav, Manish Shrivastava
2026 J jnl
ACM Comput. Surv.
Ananya Mukherjee, Manish Shrivastava
2025 conf
CODS
Prashant Kodali, Vaishnavi Shivkumar, Swarang Joshi, Monojit Choudhury, Ponnurangam Kumaraguru, Manish Shrivastava
2025 J jnl
CoRR
Prashant Kodali, Vaishnavi Shivkumar, Swarang Joshi, Monojit Choudhary, Ponnurangam Kumaraguru, Manish Shrivastava
2025 conf
NAACL (Long Papers)
Abhinav Menon, Manish Shrivastava, David Krueger, Ekdeep Singh Lubana
2025 conf
ACL (1)
Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño Cardenas, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava, Aura Cristina Udrea, Lilian Diana Awuor Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou, Saif M. Mohammad
2025 J jnl
CoRR
Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino Dário Mário António Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño Cardenas, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Andrew Piper, Alexander Panchenko, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava, Aura Cristina Udrea, Lilian Diana Awuor Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou, Saif M. Mohammad
2025 J jnl
CoRR
Aparajitha Allamraju, Maitreya Prafulla Chitale, Hiranmai Sri Adibhatla, Rahul Mishra, Manish Shrivastava
2025 conf
COLING Workshops
Likhith Asapu, Prashant Kodali, Ashna Dua, Kapil Rajesh Kavitha, Manish Shrivastava
2025 J jnl
CoRR
Ganesh Katrapati, Manish Shrivastava
2025 A* conf
ICLR
Subba Reddy Oota, Akshett Rai Jindal, Ishani Mondal, Khushbu Pahwa, Satya Sai Srinath Namburi GNVV, Manish Shrivastava, Maneesh Kumar Singh, Bapi Raju Surampudi, Manish Gupta
2025 J jnl
CoRR
Subba Reddy Oota, Akshett Jindal, Ishani Mondal, Khushbu Pahwa, Satya Sai Srinath Namburi, Manish Shrivastava, Maneesh Kumar Singh, Raju Surampudi Bapi, Manish Gupta
2025 J jnl
ACM Trans. Asian Low Resour. Lang. Inf. Process.
Prashant Kodali, Anmol Goel, Likhith Asapu, Vamshi Krishna Bonagiri, Anirudh Govil, Monojit Choudhury, Ponnurangam Kumaraguru, Manish Shrivastava
2025 J jnl
Int. J. Intell. Inf. Database Syst.
Pravin M. Tambe, Manish Shrivastava
2025 conf
NAACL (Long Papers)
Srija Mukhopadhyay, Abhishek Rajgaria, Prerana Khatiwada, Manish Shrivastava, Dan Roth, Vivek Gupta
2025 J jnl
CoRR
Sriharsh Bhyravajjula, Ujwal Narayan, Manish Shrivastava
2025 J jnl
CoRR
Manish Singh, Manish Shrivastava
2025 J jnl
CoRR
Saumitra Yadav, Manish Shrivastava
2025 conf
ACL (Findings)
Vihang Pancholi, Jainit Sushil Bafna, Tejas Anvekar, Manish Shrivastava, Vivek Gupta
2025 J jnl
CoRR
Vihang Pancholi, Jainit Sushil Bafna, Tejas Anvekar, Manish Shrivastava, Vivek Gupta
2025 B conf
COLING
Ananya Mukherjee, Saumitra Yadav, Manish Shrivastava
2024 J jnl
CoRR
Kalahasti Ganesh Srivatsa, Sabyasachi Mukhopadhyay, Ganesh Katrapati, Manish Shrivastava
2024 conf
WMT
Saumitra Yadav, Ananya Mukherjee, Manish Shrivastava
2024 J jnl
CoRR
Abhinav Menon, Manish Shrivastava, David Krueger, Ekdeep Singh Lubana
2024 J jnl
CoRR
Andrew H. Lee, Sina J. Semnani, Galo Castillo-López, Gaël de Chalendar, Monojit Choudhury, Ashna Dua, Kapil Rajesh Kavitha, Sungkyun Kim, Prashant Kodali, Ponnurangam Kumaraguru, Alexis Lombard, Mehrad Moradshahi, Gihyun Park, Nasredine Semmar, Jiwon Seo, Tianhao Shen, Manish Shrivastava, Deyi Xiong, Monica S. Lam
2024 J jnl
CoRR
Harshit Sandilya, Peehu Raj, Jainit Sushil Bafna, Srija Mukhopadhyay, Shivansh Sharma, Ellwil Sharma, Arastu Sharma, Neeta Trivedi, Manish Shrivastava, Rajesh Kumar
2024 conf
WMT
Ananya Mukherjee, Saumitra Yadav, Manish Shrivastava
2024 conf
SemEval@NAACL
Suyash Vardhan Mathur, Akshett Rai Jindal, Manish Shrivastava
2024 J jnl
CoRR
Suyash Vardhan Mathur, Akshett Rai Jindal, Manish Shrivastava
2024 J jnl
CoRR
Hiranmai Sri Adibhatla, Pavan Baswani, Manish Shrivastava
2024 J jnl
CoRR
Prashant Kodali, Anmol Goel, Likhith Asapu, Vamshi Krishna Bonagiri, Anirudh Govil, Monojit Choudhury, Manish Shrivastava, Ponnurangam Kumaraguru
2024 conf
EMNLP (Findings)
Suyash Vardhan Mathur, Jainit Sushil Bafna, Kunal Kartik, Harshita Khandelwal, Manish Shrivastava, Vivek Gupta, Mohit Bansal, Dan Roth
2024 J jnl
CoRR
Suyash Vardhan Mathur, Jainit Sushil Bafna, Kunal Kartik, Harshita Khandelwal, Manish Shrivastava, Vivek Gupta, Mohit Bansal, Dan Roth
2024 conf
SemEval@NAACL
Suyash Vardhan Mathur, Akshett Rai Jindal, Hardik Mittal, Manish Shrivastava
2024 J jnl
CoRR
Suyash Vardhan Mathur, Akshett Rai Jindal, Hardik Mittal, Manish Shrivastava
2024 conf
SemEval@NAACL
Patanjali Bhamidipati, Advaith Malladi, Manish Shrivastava, Radhika Mamidi
2024 conf
SemEval@NAACL
Jainit Sushil Bafna, Hardik Mittal, Suyash Sethia, Manish Shrivastava, Radhika Mamidi
2024 J jnl
CoRR
Jainit Sushil Bafna, Hardik Mittal, Suyash Sethia, Manish Shrivastava, Radhika Mamidi
2024 conf
SemEval@NAACL
Nedjma Ousidhoum, Shamsuddeen Hassan Muhammad, Mohamed Abdalla, Idris Abdulmumin, Ibrahim Said Ahmad, Sanchit Ahuja, Alham Fikri Aji, Vladimir Araujo, Meriem Beloucif, Christine de Kock, Oumaima Hourrane, Manish Shrivastava, Thamar Solorio, Nirmal Surange, Krishnapriya Vishnubhotla, Seid Muhie Yimam, Saif M. Mohammad
2024 J jnl
CoRR
Nedjma Ousidhoum, Shamsuddeen Hassan Muhammad, Mohamed Abdalla, Idris Abdulmumin, Ibrahim Said Ahmad, Sanchit Ahuja, Alham Fikri Aji, Vladimir Araujo, Meriem Beloucif, Christine de Kock, Oumaima Hourrane, Manish Shrivastava, Thamar Solorio, Nirmal Surange, Krishnapriya Vishnubhotla, Seid Muhie Yimam, Saif M. Mohammad
2024 conf
ACL (Findings)
Nedjma Ousidhoum, Shamsuddeen Hassan Muhammad, Mohamed Abdalla, Idris Abdulmumin, Ibrahim Said Ahmad, Sanchit Ahuja, Alham Fikri Aji, Vladimir Araujo, Abinew Ali Ayele, Pavan Baswani, Meriem Beloucif, Chris Biemann, Sofia Bourhim, Christine de Kock, Genet Shanko Dekebo, Oumaima Hourrane, Gopichand Kanumolu, Lokesh Madasu, Samuel Rutunda, Manish Shrivastava, Thamar Solorio, Nirmal Surange, Hailegnaw Getaneh Tilaye, Krishnapriya Vishnubhotla, Genta Indra Winata, Seid Muhie Yimam, Saif M. Mohammad
2024 J jnl
CoRR
Nedjma Ousidhoum, Shamsuddeen Hassan Muhammad, Mohamed Abdalla, Idris Abdulmumin, Ibrahim Said Ahmad, Sanchit Ahuja, Alham Fikri Aji, Vladimir Araujo, Abinew Ali Ayele, Pavan Baswani, Meriem Beloucif, Chris Biemann, Sofia Bourhim, Christine de Kock, Genet Shanko Dekebo, Oumaima Hourrane, Gopichand Kanumolu, Lokesh Madasu, Samuel Rutunda, Manish Shrivastava, Thamar Solorio, Nirmal Surange, Hailegnaw Getaneh Tilaye, Krishnapriya Vishnubhotla, Genta Indra Winata, Seid Muhie Yimam, Saif M. Mohammad
2024 conf
LREC/COLING
Gopichand Kanumolu, Lokesh Madasu, Nirmal Surange, Manish Shrivastava
2024 J jnl
CoRR
Gopichand Kanumolu, Lokesh Madasu, Nirmal Surange, Manish Shrivastava
2024 J jnl
CoRR
Patanjali Bhamidipati, Advaith Malladi, Manish Shrivastava, Radhika Mamidi
2024 conf
WMT
Ananya Mukherjee, Manish Shrivastava
2023 J jnl
J. Log. Lang. Inf.
Alok Debnath, Manish Shrivastava
2023 conf
SemEval@ACL
Debashish Roy, Manish Shrivastava
2023 J jnl
CoRR
Debashish Roy, Manish Shrivastava
2023 conf
COMAD/CODS
Manish Singh, Manish Shrivastava
2023 Misc conf
RANLP
Sumukh S, Abhinav Appidi Reddy, Manish Shrivastava
2023 C conf
PACLIC
Hiranmai Sri Adibhatla, Pavan Baswani, Manish Shrivastava
2023 conf
WMT
Ananya Mukherjee, Manish Shrivastava
2023 J jnl
CoRR
Ashok Urlana, Sahil Manoj Bhatt, Nirmal Surange, Manish Shrivastava
2023 conf
SemEval@ACL
Pavan Baswani, Hiranmai Sri Adibhatla, Manish Shrivastava
2023 conf
Eval4NLP
Pavan Baswani, Ananya Mukherjee, Manish Shrivastava
2023 conf
WMT
Ananya Mukherjee, Manish Shrivastava
2023 C conf
PACLIC
Lokesh Madasu, Gopichand Kanumolu, Nirmal Surange, Manish Shrivastava
2023 J jnl
CoRR
Lokesh Madasu, Gopichand Kanumolu, Nirmal Surange, Manish Shrivastava
2023 conf
ICMI Companion
Mounika Kanakanti, Shantanu Singh, Manish Shrivastava
2023 conf
EMNLP (Findings)
Ashok Urlana, Pinzhen Chen, Zheng Zhao, Shay B. Cohen, Manish Shrivastava, Barry Haddow
2023 J jnl
CoRR
Ashok Urlana, Pinzhen Chen, Zheng Zhao, Shay B. Cohen, Manish Shrivastava, Barry Haddow
2023 conf
WASSA@ACL
Bhaskara Hanuma Vedula, Prashant Kodali, Manish Shrivastava, Ponnurangam Kumaraguru
2023 C conf
PACLIC
Shubhangi Dutta, Manish Shrivastava, Prabhakar Bhimalapuram
2023 conf
ECIR (1)
Sahil Manoj Bhatt, Sahaj Agarwal, Omkar Gurjar, Manish Gupta, Manish Shrivastava
2023 J jnl
CoRR
Gopichand Kanumolu, Lokesh Madasu, Pavan Baswani, Ananya Mukherjee, Manish Shrivastava
2023 conf
ACL (Findings)
Mehrad Moradshahi, Tianhao Shen, Kalika Bali, Monojit Choudhury, Gaël de Chalendar, Anmol Goel, Sungkyun Kim, Prashant Kodali, Ponnurangam Kumaraguru, Nasredine Semmar, Sina J. Semnani, Jiwon Seo, Vivek Seshadri, Manish Shrivastava, Michael Sun, Aditya Yadavalli, Chaobin You, Deyi Xiong, Monica S. Lam
2023 J jnl
CoRR
Mehrad Moradshahi, Tianhao Shen, Kalika Bali, Monojit Choudhury, Gaël de Chalendar, Anmol Goel, Sungkyun Kim, Prashant Kodali, Ponnurangam Kumaraguru, Nasredine Semmar, Sina J. Semnani, Jiwon Seo, Vivek Seshadri, Manish Shrivastava, Michael Sun, Aditya Yadavalli, Chaobin You, Deyi Xiong, Monica S. Lam
2022 conf
W-NUT@COLING
Sumukh S, Manish Shrivastava
2022 conf
NAACL-HLT
Chaitanya Agarwal, Vivek Gupta, Anoop Kunchukuttan, Manish Shrivastava
2022 B conf
COLING
Poojitha Nandigam, Nikhil Rayaprolu, Manish Shrivastava
2022 J jnl
CoRR
Poojitha Nandigam, Nikhil Rayaprolu, Manish Shrivastava
2022 A* conf
EMNLP
Puneet Mathur, Gautam Kunapuli, Riyaz A. Bhat, Manish Shrivastava, Dinesh Manocha, Maneesh Singh
2022 conf
ICON
Souvik Banerjee, Bamdev Mishra, Pratik Jawanpuria, Manish Shrivastava
2022 J jnl
CoRR
Souvik Banerjee, Bamdev Mishra, Pratik Jawanpuria, Manish Shrivastava
2022 B conf
LREC
Prashant Kodali, Akshala Bhatnagar, Naman Ahuja, Manish Shrivastava, Ponnurangam Kumaraguru
2022 J jnl
CoRR
Prashant Kodali, Akshala Bhatnagar, Naman Ahuja, Manish Shrivastava, Ponnurangam Kumaraguru
2022 conf
FIRE (Working Notes)
Ashok Urlana, Sahil Manoj Bhatt, Nirmal Surange, Manish Shrivastava
2022 J jnl
Trans. Assoc. Comput. Linguistics
Vivek Gupta, Riyaz A. Bhat, Atreya Ghosal, Manish Shrivastava, Maneesh Kumar Singh, Vivek Srikumar
2022 conf
CASE@EMNLP
Hiranmai Sri Adibhatla, Manish Shrivastava
2022 conf
SDP@COLING
Ashok Urlana, Nirmal Surange, Manish Shrivastava
2022 conf
EMNLP (Findings)
Aashna Jena, Vivek Gupta, Manish Shrivastava, Julian Martin Eisenschlos
2022 J jnl
CoRR
Aashna Jena, Vivek Gupta, Manish Shrivastava, Julian Martin Eisenschlos
2022 conf
Text2Story@ECIR
Sriharsh Bhyravajjula, Ujwal Narayan, Manish Shrivastava
2022 conf
ICON
Poojitha Nandigam, Abhinav Appidi Reddy, Manish Shrivastava
2022 J jnl
CoRR
Prashant Kodali, Tanmay Sachan, Akshay Goindani, Anmol Goel, Naman Ahuja, Manish Shrivastava, Ponnurangam Kumaraguru
2022 conf
WMT
Ananya Mukherjee, Manish Shrivastava
2022 conf
ICON
Hiranmai Sri Adibhatla, Manish Shrivastava
2022 conf
ACL (Findings)
Prashant Kodali, Anmol Goel, Monojit Choudhury, Manish Shrivastava, Ponnurangam Kumaraguru
2022 B conf
LREC
Ashok Urlana, Nirmal Surange, Pawan Baswani, Priyanka Ravva, Manish Shrivastava
2022 conf
SemEval@NAACL
Sahil Manoj Bhatt, Manish Shrivastava
2022 conf
ACL (student)
Roopal Vaid, Kartikey Pant, Manish Shrivastava
2022 conf
WMT
Ananya Mukherjee, Manish Shrivastava
2021 conf
WebSci
Mohit Chandra, Dheeraj Reddy Pailla, Himanshu Bhatia, AadilMehdi J. Sanchawala, Manish Gupta, Manish Shrivastava, Ponnurangam Kumaraguru
2021 J jnl
CoRR
Mohit Chandra, Dheeraj Reddy Pailla, Himanshu Bhatia, AadilMehdi J. Sanchawala, Manish Gupta, Manish Shrivastava, Ponnurangam Kumaraguru
2021 Misc conf
RANLP
Akshay Goindani, Manish Shrivastava
2021 J jnl
CoRR
Akshay Goindani, Manish Shrivastava
2021 conf
WMT@EMNLP
Saumitra Yadav, Manish Shrivastava
2021 J jnl
CoRR
Aditya Kadam, Anmol Goel, Jivitesh Jain, Jushaan Singh Kalra, Mallika Subramanian, Manvith Reddy, Prashant Kodali, T. H. Arjun, Manish Shrivastava, Ponnurangam Kumaraguru
2021 conf
FIRE (Working Notes)
Aditya Kadam, Anmol Goel, Jivitesh Jain, Jushaan Singh Kalra, Mallika Subramanian, Manvith Reddy, Prashant Kodali, T. H. Arjun, Manish Shrivastava, Ponnurangam Kumaraguru
2021 conf
CALCS@NAACL
Devansh Gautam, Prashant Kodali, Kshitij Gupta, Anmol Goel, Manish Shrivastava, Ponnurangam Kumaraguru
2021 conf
COMAD/CODS
Manish Singh, Manish Shrivastava
2021 J jnl
CoRR
Vivek Gupta, Riyaz A. Bhat, Atreya Ghosal, Manish Shrivastava, Maneesh Singh, Vivek Srikumar
2021 B conf
SIGDIAL
Konigari Rachna, Saurabh Ramola, Vijay Vardhan Alluri, Manish Shrivastava
2021 conf
CALCS@NAACL
Devansh Gautam, Kshitij Gupta, Manish Shrivastava
2021 conf
SemEval@ACL/IJCNLP
Devansh Gautam, Kshitij Gupta, Manish Shrivastava
2021 J jnl
CoRR
Devansh Gautam, Kshitij Gupta, Manish Shrivastava
2020 conf
TRAC@LREC
Arjit Srivastava, Avijit Vajpayee, Syed Sarfaraz Akhtar, Naman Jain, Vinay Singh, Manish Shrivastava
2020 conf
ACL (student)
Sneha Nallani, Manish Shrivastava, Dipti Misra Sharma
2020 conf
WMT@EMNLP
Saumitra Yadav, Manish Shrivastava
2020 conf
COMAD/CODS
Priyanka Ravva, Ashok Urlana, Manish Shrivastava
2020 B conf
COLING
Mohit Chandra, Ashwin Pathak, Eesha Dutta, Paryul Jain, Manish Gupta, Manish Shrivastava, Ponnurangam Kumaraguru
2020 J jnl
CoRR
Mohit Chandra, Ashwin Pathak, Eesha Dutta, Paryul Jain, Manish Gupta, Manish Shrivastava, Ponnurangam Kumaraguru
2020 conf
TDS
Vaishali Pal, Manish Shrivastava, Laurent Besacier
2020 J jnl
CoRR
Vaishali Pal, Manish Shrivastava, Laurent Besacier
2020 conf
ICON
Appidi Abhinav Reddy, Vamshi Krishna Srirangam, Darsi Suhas, Manish Shrivastava
2020 B conf
COLING
Appidi Abhinav Reddy, Vamshi Krishna Srirangam, Darsi Suhas, Manish Shrivastava
2020 conf
TDS
Anirudh Dahiya, Manish Shrivastava, Dipti Misra Sharma
2020 B conf
CoNLL
Payal Khullar, Arghya Bhattacharya, Manish Shrivastava
2020 conf
ICON
Chaitanya Sai Alaparthi, Manish Shrivastava
2020 B conf
DSAA
Ananya Mukherjee, Hema Ala, Manish Shrivastava, Dipti Misra Sharma
2020 A conf
INTERSPEECH
Vaishali Pal, Fabien Guillot, Manish Shrivastava, Jean-Michel Renders, Laurent Besacier
2020 B conf
LREC
Payal Khullar, Kushal Majmundar, Manish Shrivastava
2020 conf
ECIR (2)
Muthusamy Chelliah, Manish Shrivastava, Jaidam Ram Tej
2020 conf
ACL (student)
Chanakya Malireddy, Tirth Maniar, Manish Shrivastava
2020 conf
SemEval@COLING
Sravani Boinepelli, Manish Shrivastava, Vasudeva Varma
2020 conf
WWW (Companion Volume)
Alok Debnath, Nikhil Pinnaparaju, Manish Shrivastava, Vasudeva Varma, Isabelle Augenstein
2020 conf
ISMIR
Aayush Surana, Yash Goyal, Manish Shrivastava, Suvi Saarikallio, Vinoo Alluri
2020 J jnl
CoRR
Aayush Surana, Yash Goyal, Manish Shrivastava, Suvi Saarikallio, Vinoo Alluri
2020 conf
ICON
Ravindra Nittala, Manish Shrivastava
2020 conf
RepL4NLP@ACL
Siddharth Bhat, Alok Debnath, Souvik Banerjee, Manish Shrivastava
2019 conf
NAACL-HLT (Student Research Workshop)
Alok Debnath, Manish Shrivastava
2019 conf
ACL (2)
Vamshi Krishna Srirangam, Appidi Abhinav Reddy, Vinay Singh, Manish Shrivastava
2019 A* conf
IJCAI
Anirudh Dahiya, Neeraj Battan, Manish Shrivastava, Dipti Misra Sharma
2019 J jnl
CoRR
Anirudh Dahiya, Neeraj Battan, Manish Shrivastava, Dipti Misra Sharma
2019 conf
ACL (2)
Yash Kumar Lal, Vaibhav Kumar, Mrinal Dhar, Manish Shrivastava, Philipp Koehn
2019 conf
SemEval@NAACL-HLT
Vijayasaradhi Indurthi, Bakhtiyar Syed, Manish Shrivastava, Nikhil Chakravartula, Manish Gupta, Vasudeva Varma
2019 conf
SemEval@NAACL-HLT
Vijayasaradhi Indurthi, Bakhtiyar Syed, Manish Shrivastava, Manish Gupta, Vasudeva Varma
2019 conf
SemEval@NAACL-HLT
Bakhtiyar Syed, Vijayasaradhi Indurthi, Manish Shrivastava, Manish Gupta, Vasudeva Varma
2019 conf
ECIR (2)
Bakhtiyar Syed, Vijayasaradhi Indurthi, Manish Gupta, Manish Shrivastava, Vasudeva Varma
2019 conf
W-NUT@EMNLP
Vinayak Athavale, Aayush Naik, Rajas Vanjape, Manish Shrivastava
2019 J jnl
CoRR
Vinayak Athavale, Aayush Naik, Rajas Vanjape, Manish Shrivastava
2019 J jnl
CoRR
Ratish Puduppully, Yue Zhang, Manish Shrivastava
2019 conf
ICSC
Krishna Rohit Sakala Venkata, Manish Shrivastava
2019 Misc conf
RANLP
Payal Khullar, Allen Antony, Manish Shrivastava
2018 conf
ICON
Faraz Faruqi, Manish Shrivastava
2018 J jnl
CoRR
Sahil Swami, Ankush Khandelwal, Vinay Singh, Syed Sarfaraz Akhtar, Manish Shrivastava
2018 conf
EMSASW@ESWC
Deepanshu Vijay, Aditya Bohra, Vinay Singh, Syed Sarfaraz Akhtar, Manish Shrivastava
2018 conf
PEOPLES@NAACL-HTL
Aditya Bohra, Deepanshu Vijay, Vinay Singh, Syed Sarfaraz Akhtar, Manish Shrivastava
2018 conf
ALW
Vinay Singh, Aman Varshney, Syed Sarfaraz Akhtar, Deepanshu Vijay, Manish Shrivastava
2018 J jnl
CoRR
Sahil Swami, Ankush Khandelwal, Vinay Singh, Syed Sarfaraz Akhtar, Manish Shrivastava
2018 conf
CICLing (1)
Rajat Singh, Nurendra Choudhary, Manish Shrivastava
2018 J jnl
CoRR
Rajat Singh, Nurendra Choudhary, Manish Shrivastava
2018 conf
ACL (3)
Payal Khullar, Konigari Rachna, Mukul Hase, Manish Shrivastava
2018 C conf
PACLIC
Pranav Dhakras, Manish Shrivastava
2018 conf
CICLing (2)
Nurendra Choudhary, Rajat Singh, Ishita Bindlish, Manish Shrivastava
2018 J jnl
CoRR
Nurendra Choudhary, Rajat Singh, Ishita Bindlish, Manish Shrivastava
2018 conf
NAACL-HLT (Student Research Workshop)
Deepanshu Vijay, Aditya Bohra, Vinay Singh, Syed Sarfaraz Akhtar, Manish Shrivastava
2018 J jnl
CoRR
Nurendra Choudhary, Rajat Singh, Manish Shrivastava
2018 conf
ICSE (Companion Volume)
Amar Budhiraja, Kartik Dutta, Raghu Reddy, Manish Shrivastava
2018 conf
ICON
Aishwary Gupta, Akshay Pawale, Manish Shrivastava
2018 conf
TRAC@COLING 2018
Sanjana Sharma, Saksham Agrawal, Manish Shrivastava
2018 J jnl
CoRR
Sanjana Sharma, Saksham Agrawal, Manish Shrivastava
2018 conf
CICLing (2)
Nurendra Choudhary, Rajat Singh, Ishita Bindlish, Manish Shrivastava
2018 J jnl
CoRR
Nurendra Choudhary, Rajat Singh, Ishita Bindlish, Manish Shrivastava
2018 conf
ACL (3)
Nikhilesh Bhatnagar, Manish Shrivastava, Radhika Mamidi
2018 J jnl
CoRR
Ankush Khandelwal, Sahil Swami, Syed Sarfaraz Akhtar, Manish Shrivastava
2018 J jnl
Computación y Sistemas
Ankush Khandelwal, Sahil Swami, Syed Sarfaraz Akhtar, Manish Shrivastava
2018 B conf
LREC
Ankush Khandelwal, Sahil Swami, Syed Sarfaraz Akhtar, Manish Shrivastava
2018 J jnl
CoRR
Ankush Khandelwal, Sahil Swami, Syed Sarfaraz Akhtar, Manish Shrivastava
2018 conf
ICSE (Companion Volume)
Amar Budhiraja, Raghu Reddy, Manish Shrivastava
2018 conf
NEWS@ACL
Vinay Singh, Deepanshu Vijay, Syed Sarfaraz Akhtar, Manish Shrivastava
2018 conf
CICLing (2)
Nurendra Choudhary, Rajat Singh, Ishita Bindlish, Manish Shrivastava
2018 J jnl
CoRR
Rajat Singh, Nurendra Choudhary, Ishita Bindlish, Manish Shrivastava
2018 conf
FEVER@EMNLP
Vishal Gupta, Manoj Kumar Chinnakotla, Manish Shrivastava
2018 J jnl
CoRR
Vaibhav Kumar, Mrinal Dhar, Dhruv Khattar, Yash Kumar Lal, Abhimanshu Mishra, Manish Shrivastava, Vasudeva Varma
2018 conf
CICLing (2)
Nurendra Choudhary, Rajat Singh, Ishita Bindlish, Manish Shrivastava
2018 J jnl
CoRR
Nurendra Choudhary, Rajat Singh, Ishita Bindlish, Manish Shrivastava
2018 C conf
PACLIC
Prathyusha Danda, Brij Mohan Lal Srivastava, Manish Shrivastava
2018 Misc conf
ICTIR
Amar Budhiraja, Kartik Dutta, Manish Shrivastava, Raghu Reddy
2018 conf
CodeSwitch@ACL
Vishal Gupta, Manoj Kumar Chinnakotla, Manish Shrivastava
2018 conf
ICCSA (6)
Rashid Ahmad, Priyank Gupta, Nagaraju Vuppala, Sanket Kumar Pathak, Ashutosh Kumar, Gagan Soni, Sravan Kumar, Manish Shrivastava, Avinash K. Singh, Arbind K. Gangwar, Pawan Kumar, Mukul K. Sinha
2018 B conf
COLING
Nurendra Choudhary, Rajat Singh, Vijjini Anvesh Rao, Manish Shrivastava
2018 conf
NAACL-HLT
Irshad Ahmad Bhat, Riyaz Ahmad Bhat, Manish Shrivastava, Dipti Misra Sharma
2018 J jnl
CoRR
Irshad Ahmad Bhat, Riyaz Ahmad Bhat, Manish Shrivastava, Dipti Misra Sharma
2017 J jnl
CoRR
Syed Sarfaraz Akhtar, Arihant Gupta, Avijit Vajpayee, Arjit Srivastava, Madan Gopal Jhawar, Manish Shrivastava
2017 conf
ICON
Vijay Prakash Dwivedi, Manish Shrivastava
2017 conf
IberEval@SEPLN
Ankush Khandelwal, Sahil Swami, Syed Sarfaraz Akhtar, Manish Shrivastava
2017 conf
MIKE
Vishnu Vidyadhara Raju Vegesna, Krishna Gurugubelli, Hari Krishna Vydana, Bhargav Pulugandla, Manish Shrivastava, Anil Kumar Vuppala
2017 conf
IJCNLP (System Demonstrations)
Purvanshi Mehta, Pruthwik Mishra, Vinayak Athavale, Manish Shrivastava, Dipti Misra Sharma
2017 conf
ICON
Prathyusha Danda, Prathyusha Jwalapuram, Manish Shrivastava
2017 A* conf
EMNLP
Arihant Gupta, Syed Sarfaraz Akhtar, Avijit Vajpayee, Arjit Srivastava, Madan Gopal Jhanwar, Manish Shrivastava
2017 conf
ICCSA (7)
Priyank Gupta, Rashid Ahmad, Manish Shrivastava, Pawan Kumar, Mukul K. Sinha
2017 conf
IJCNLP(2)
Prakhar Pandey, Vikram Pudi, Manish Shrivastava
2017 conf
EACL (2)
Irshad Ahmad Bhat, Riyaz Ahmad Bhat, Manish Shrivastava, Dipti Misra Sharma
2017 J jnl
CoRR
Irshad Ahmad Bhat, Riyaz Ahmad Bhat, Manish Shrivastava, Dipti Misra Sharma
2017 conf
IberEval@SEPLN
Sahil Swami, Ankush Khandelwal, Manish Shrivastava, Syed Sarfaraz Akhtar
2017 J jnl
CoRR
Nausheen Fatma, Manoj Kumar Chinnakotla, Manish Shrivastava
2017 conf
IC3
Harika Abburi, K. N. R. K. Raju Alluri, Anil Kumar Vuppala, Manish Shrivastava, Suryakanth V. Gangashetty
2017 conf
MIKE
Harika Abburi, Rajendra Prasath, Manish Shrivastava, Suryakanth V. Gangashetty
2017 B conf
IJCNN
Brij Mohan Lal Srivastava, Hari Krishna Vydana, Anil Kumar Vuppala, Manish Shrivastava
2017 A* conf
AAAI
Nausheen Fatma, Manoj Kumar Chinnakotla, Manish Shrivastava
2017 conf
EACL (1)
Ratish Puduppully, Yue Zhang, Manish Shrivastava
2017 J jnl
CoRR
Syed Sarfaraz Akhtar, Arihant Gupta, Avijit Vajpayee, Arjit Srivastava, Manish Shrivastava
2017 conf
CLEF
Khyathi Raghavi Chandu, Manoj Kumar Chinnakotla, Alan W. Black, Manish Shrivastava
2017 conf
LAW@ACL
Syed Sarfaraz Akhtar, Arihant Gupta, Avijit Vajpayee, Arjit Srivastava, Manish Shrivastava
2016 conf
SLSP
Brij Mohan Lal Srivastava, Manish Shrivastava
2016 conf
FIRE (Working Notes)
Irshad Ahmad Bhat, Manish Shrivastava, Riyaz Ahmad Bhat
2016 J jnl
CoRR
Sai Praneeth Suggu, Kushwanth N. Goutham, Manoj Kumar Chinnakotla, Manish Shrivastava
2016 B conf
COLING
Sai Praneeth Suggu, Kushwanth Naga Goutham, Manoj Kumar Chinnakotla, Manish Shrivastava
2016 conf
HLT-NAACL Demos
Sharada Prasanna Mohanty, Nehal J. Wani, Manish Shrivastava, Dipti Misra Sharma
2016 conf
PAKDD (2)
Arpita Das, Manish Shrivastava, Manoj Kumar Chinnakotla
2016 conf
MIKE
Harika Abburi, Rajendra Prasath, Manish Shrivastava, Suryakanth V. Gangashetty
2016 conf
HLT-NAACL
Arnav Sharma, Sakshi Gupta, Raveesh Motlani, Piyush Bansal, Manish Shrivastava, Radhika Mamidi, Dipti Misra Sharma
2016 J jnl
CoRR
Arnav Sharma, Sakshi Gupta, Raveesh Motlani, Piyush Bansal, Manish Shrivastava, Radhika Mamidi, Dipti Misra Sharma
2016 conf
ACL (1)
Arpita Das, Harish Yenala, Manoj Kumar Chinnakotla, Manish Shrivastava
2016 conf
ICON
Vinayak Athavale, Shreenivas Bharadwaj, Monik Pamecha, Ameya Prabhu, Manish Shrivastava
2016 J jnl
CoRR
Vinayak Athavale, Shreenivas Bharadwaj, Monik Pamecha, Ameya Prabhu, Manish Shrivastava
2016 B conf
COLING
Aditya Joshi, Ameya Prabhu, Manish Shrivastava, Vasudeva Varma
2016 J jnl
CoRR
Ameya Prabhu, Aditya Joshi, Manish Shrivastava, Vasudeva Varma
2016 conf
HLT-NAACL
Ratish Puduppully, Yue Zhang, Manish Shrivastava
2016 conf
ICON
Prathyusha Danda, Brij Mohan Lal Srivastava, Manish Shrivastava
2015 conf
WWW (Companion Volume)
Khyathi Chandu Raghavi, Manoj Kumar Chinnakotla, Manish Shrivastava
2015 J jnl
CoRR
Brij Mohan Lal Srivastava, Hari Krishna Vydana, Anil Kumar Vuppala, Manish Shrivastava
2015 conf
CLEF (Working Notes)
Harish Yenala, Avinash Kamineni, Manish Shrivastava, Manoj Kumar Chinnakotla
2015 conf
CLEF (Working Notes)
Avinash Kamineni, Nausheen Fatma, Arpita Das, Manish Shrivastava, Manoj Kumar Chinnakotla
2014 C conf
GWC
Diptesh Kanojia, Pushpak Bhattacharyya, Raj Dabre, Siddhartha Gunti, Manish Shrivastava
2014 conf
FIRE
Irshad Ahmad Bhat, Vandan Mujadia, Aniruddha Tammewar, Riyaz Ahmad Bhat, Manish Shrivastava
2014 conf
ICON
Diptesh Kanojia, Manish Shrivastava, Raj Dabre, Pushpak Bhattacharyya
2013 conf
WOCN
Neha Gupta, Rajeev Kumar Singh, Manish Shrivastava
2006 A* conf
ACL
Smriti Singh, Kuhoo Gupta, Manish Shrivastava, Pushpak Bhattacharyya
redb/workers.py
← Index redb/workers.py python
"""
Worker and file processing functions for REDB.

Contains all functions that process binary files, handle archive extraction,
run extractors, and manage worker processes for parallel processing.
"""
import hashlib
import sys
import logging
from typing import List, Type, Optional
import os
import gc
import signal
import psutil
import time
import tempfile
import magic
from magika import Magika

import py7zr
import pyzipper

from redb import settings
from redb.extractors.basicproperties import BasicPropertiesExtractor
from redb.extractors.detectiteasy import DIEExtractor
from redb.extractors.capa import CAPAExtractor
from redb.extractors.malcontent import MalcontentExtractor
from redb.extractors.hashes import HashExtractor
from redb.extractors.extractor import Extractor
from redb.extractors.yara import YaraExtractor
from redb.extractors.database_exporters import ElasticsearchExporter, ClickHouseExporter, PrintExporter
from redb.extractor_registry import get_filetype_modules, get_extractor_class

from redb.logging_utils import ImportResult, setup_direct_logger, setup_logger
from redb.queries import is_in_db, is_in_code_db
from redb.s3_utils import download_s3_object, extract_fat_slices



def get_module_by_name(module_name: str) -> Type[Extractor]:
    """Get module class by its name."""
    # Common extractors (always loaded, lightweight)
    common_map = {
        "BasicPropertiesExtractor": BasicPropertiesExtractor,
        "HashExtractor": HashExtractor,
        "DIEExtractor": DIEExtractor,
        "CAPAExtractor": CAPAExtractor,
        "MalcontentExtractor": MalcontentExtractor,
        "YaraExtractor": YaraExtractor,
    }
    if module_name in common_map:
        return common_map[module_name]
    # File-type-specific extractors (loaded on demand via registry)
    return get_extractor_class(module_name)


def filter_selected_modules(
    selected_modules: List[str], available_modules: List[Type[Extractor]]
) -> List[Type[Extractor]]:
    """Filter available modules based on selected module names."""
    if not selected_modules or "all" in selected_modules:
        return available_modules

    return [mod for mod in available_modules if mod.__name__ in selected_modules]


def _is_packed(sha256, index_prefix, logger):
    """
    Check if a file is packed by querying ClickHouse for DIE packer information.

    Args:
        sha256: File's SHA256 hash
        index_prefix: Elastic index prefix
        logger: Logger instance

    Returns:
        bool or None:
            - True if file exists and is packed
            - False if file exists but is not packed
            - None if file doesn't exist in the index
    """
    try:
        table = f"{index_prefix}_die_mv"

        client = settings.create_clickhouse_client()
        packer_query = f"""
            SELECT count(*) FROM {table}
            WHERE sha256 = '{sha256}'
            AND packer IS NOT NULL
        """
        packer_result = client.query(packer_query).result_rows[0][0]

        if packer_result > 0:
            logger.debug(f"File {sha256} is packed")
            return True
        else:
            logger.debug(f"File {sha256} is not packed")
            return False
    except Exception as e:
        logger.info(f"File {sha256} packed status UNKNOWN")
        logger.error(f"Error checking if file is packed: {str(e)}")
        return None


def is_binary_file(file):
    try:
        with open(file, "tr") as check_file:
            check_file.read()
        return False
    except:
        return True


def _is_supported_format(filepath, logger):
    """Check if a non-binary file has a format listed in SUPPORTED_FORMATS (.env).

    Called only when is_binary_file returns False (i.e. the file is text-based).
    Uses Magika to detect the format, then checks against SUPPORTED_FORMATS.
    """
    try:
        from redb.queries import get_supported_formats

        with open(filepath, "rb") as f:
            data = f.read()
        filetype = Magika().identify_bytes(data).output.label
        if filetype in get_supported_formats():
            logger.debug(f"Detected supported text format: {filetype}")
            return True
    except Exception as e:
        logger.debug(f"Error detecting format for {filepath}: {e}")
    return False



def check_dotnet(path_to_parse, logger):
    try:
        import pefile
        file_type = magic.from_file(path_to_parse)
        if ".Net" in file_type:
            return True
        pe = pefile.PE(path_to_parse)
        for entry in pe.OPTIONAL_HEADER.DATA_DIRECTORY:
            if entry.name == "IMAGE_DIRECTORY_ENTRY_COM_DESCRIPTOR" and entry.Size > 0:
                return True
        return False
    except AttributeError as e:
        logger.error(
            f"AttributeError error dotnet file {path_to_parse} Full error : {e}"
        )
        return False


def check_high_swap():
    """Check if swap usage is critically high and take action if needed"""
    try:
        swap_usage = psutil.swap_memory().percent
        if swap_usage > 90:
            print(f"[WARNING] Critical swap usage detected: {swap_usage}%")

            # Force immediate garbage collection
            gc.collect(2)  # Full collection with generation 2
            return True
    except:
        pass
    return False


def process_s3_file(
    s3_bucket: str,
    s3_key: str,
    temp_dir: str,
    decompile: bool,
    index_prefix: str,
    log_file: str,
    file_number: int,
    total_files: int,
    selected_modules=None,
    dry_run=False,
    yara_scan=False,
    with_yara=False,
    force=False,
    decompile_modules=None,
    first_seen=None,
):
    """
    Download and process a file from S3
    """
    filename = os.path.basename(s3_key)
    extra = {"filenameinfo": filename}
    worker_pid = os.getpid()
    logger = setup_direct_logger(log_file, filename, worker_pid)

    logger.info(f"Processing S3 file {file_number}/{total_files}: {s3_key}")

    try:
        # Download file from S3
        local_path = download_s3_object(s3_bucket, s3_key, temp_dir, logger)
        if not local_path:
            logger.error(f"Failed to download {s3_key}")
            return ImportResult.FAILED, None

        # Process the downloaded file using existing logic
        return _process_file_internal(
            local_path, decompile, index_prefix, logger, selected_modules, dry_run, yara_scan, with_yara, force,
            decompile_modules=decompile_modules,
            first_seen=first_seen,
        )
    except Exception as e:
        logger.error(f"Error processing S3 file: {str(e)}")
        return ImportResult.FAILED, None


def process_file(
    filepath,
    decompile,
    index_prefix,
    log_file,
    file_number,
    total_files,
    selected_modules=None,
    dry_run=False,
    yara_scan=False,
    with_yara=False,
    force=False,
    decompile_modules=None,
):
    filename = os.path.basename(filepath)
    worker_pid = os.getpid()

    # Add filename info to the logger's extra info
    extra = {"filenameinfo": filename}

    logger = setup_direct_logger(log_file, filename, worker_pid)
    logger.info(f"Processing file {file_number}/{total_files}: {filepath}")

    try:
        return _process_file_internal(
            filepath, decompile, index_prefix, logger, selected_modules, dry_run, yara_scan, with_yara, force,
            decompile_modules=decompile_modules,
        )
    except Exception as e:
        logger.error(f"Error processing file: {str(e)}")
        return ImportResult.FAILED, None


def _process_file_internal(
    filepath, decompile, index_prefix, logger, selected_modules=None, dry_run=False, yara_scan=False, with_yara=False, force=False,
    decompile_modules=None, first_seen=None,
):
    _, ext = os.path.splitext(filepath.lower())
    if ext == ".zip":
        return process_zip_file(
            filepath, decompile, index_prefix, logger, selected_modules, dry_run, yara_scan, with_yara, force,
            decompile_modules=decompile_modules, first_seen=first_seen)
    elif ext == ".7z":
        return process_7zip_file(
            filepath, decompile, index_prefix, logger, selected_modules, dry_run, yara_scan, with_yara, force,
            decompile_modules=decompile_modules, first_seen=first_seen)
    elif is_binary_file(filepath) or _is_supported_format(filepath, logger):
        return process_binary_file(
            filepath, decompile, index_prefix, logger, selected_modules, dry_run, yara_scan, with_yara, force,
            decompile_modules=decompile_modules, first_seen=first_seen)
    else:
        logger.info(f"Skipping non-binary file: {filepath}")
        return ImportResult.SKIPPED, None


def process_zip_file(filepath, decompile, index_prefix, logger, selected_modules, dry_run=False, yara_scan=False, with_yara=False, force=False, decompile_modules=None, first_seen=None):
    logger.info(f"Processing zip file: {filepath}")
    with tempfile.TemporaryDirectory() as temp_dir:
        extracted_zip_path = None

        # Try standard zipfile first (ZipCrypto encryption - used by zipencrypt)
        try:
            import zipfile
            with zipfile.ZipFile(filepath, 'r') as zf:
                filename = zf.namelist()[0]
                # Must pass pwd directly to extractall(), setpassword() alone doesn't work
                zf.extractall(temp_dir, pwd=b"infected")
                extracted_zip_path = os.path.join(temp_dir, filename)
                logger.debug(f"Extracted with standard zipfile (ZipCrypto): {extracted_zip_path}")
        except Exception as e:
            logger.debug(f"Standard zipfile extraction failed: {str(e)}, trying AES...")

        # If standard zipfile failed, try pyzipper for AES encryption
        if not extracted_zip_path or not os.path.exists(extracted_zip_path):
            try:
                with pyzipper.AESZipFile(filepath) as zf:
                    zf.pwd = b"infected"
                    filename = zf.namelist()[0]
                    zf.extractall(temp_dir)
                    extracted_zip_path = os.path.join(temp_dir, filename)
                    logger.debug(f"Extracted with pyzipper (AES): {extracted_zip_path}")
            except Exception as e:
                logger.error(f"Error processing zip file {filepath}: {str(e)}")
                return ImportResult.FAILED, None

        # Process the extracted binary
        if extracted_zip_path and os.path.exists(extracted_zip_path):
            return process_binary_file(
                extracted_zip_path,
                decompile,
                index_prefix,
                logger,
                selected_modules,
                dry_run,
                yara_scan,
                with_yara,
                force,
                decompile_modules=decompile_modules,
                first_seen=first_seen)
        else:
            logger.error(f"Failed to extract zip file {filepath}")
            return ImportResult.FAILED, None


def process_7zip_file(filepath, decompile, index_prefix, logger, selected_modules, dry_run=False, yara_scan=False, with_yara=False, force=False, decompile_modules=None, first_seen=None):
    logger.info(f"Processing 7zip file: {filepath}")
    with tempfile.TemporaryDirectory() as temp_dir:
        try:
            with py7zr.SevenZipFile(filepath, mode="r", password="infected") as z:
                z.extractall(path=temp_dir)
                logger.debug(f"Extracted 7zip to: {temp_dir}")
                for root, _, files in os.walk(temp_dir):
                    for file in files:
                        extracted_file_path = os.path.join(root, file)
                        logger.debug(
                            f"Processing extracted file: {extracted_file_path}"
                        )
                        return process_binary_file(
                            extracted_file_path,
                            decompile,
                            index_prefix,
                            logger,
                            selected_modules,
                            dry_run,
                            yara_scan,
                            with_yara,
                            force,
                            decompile_modules=decompile_modules,
                            first_seen=first_seen)
        except Exception as e:
            logger.error(f"Error processing 7zip file {filepath}: {str(e)}")
            return ImportResult.FAILED, None


def process_binary_file(
    filepath, decompile, index_prefix, logger, selected_modules=None, dry_run=False, yara_scan=False, with_yara=False, force=False,
    decompile_modules=None, first_seen=None):
    """
    Process a binary file with selected modules.

    Args:
        filepath: Path to the file
        decompile: Boolean flag for decompilation
        index_prefix: Elastic index prefix
        logger: Logger instance
        selected_modules: List of module names to run, or None/'all' for all modules
        dry_run: Boolean flag for dry-run mode
        yara_scan: Boolean flag for YARA scanning only mode
        with_yara: Boolean flag for combined features + YARA mode
    """
    start_time = time.time()

    with open(filepath, "rb") as f:
        data = f.read()
    filetype = Magika().identify_bytes(data).output.label
    sha256 = hashlib.sha256(data).hexdigest()

    # Initialize exporters
    exporters = []

    clickhouse_client = None
    if dry_run:
        # Use PrintExporter for dry-run mode
        print_exporter = PrintExporter(logger, index_prefix)
        exporters.append(print_exporter)
        logger.info("DRY-RUN mode: Results will be printed instead of uploaded to database")
    elif settings.CLICKHOUSE_HOST:  # Check if ClickHouse is configured
        try:
            clickhouse_client = settings.get_clickhouse_client()
            ch_exporter = ClickHouseExporter(logger, index_prefix, client=clickhouse_client)
            exporters.append(ch_exporter)
            logger.debug("ClickHouse exporter initialized successfully")
        except Exception as e:
            logger.error(f"Failed to create ClickHouse client: {e}")
            return ImportResult.FAILED, None
    else:
        logger.error("No exporters configured - neither dry-run mode nor ClickHouse available")
        return ImportResult.FAILED, None

    # Check if file already exists in database (deduplication)
    # Skip this check for decompile mode - it has its own is_in_code_db checks
    # Skip this check when force is True (--force flag or specific --modules)
    # Skip this check for yara_scan mode - YARA dedup is handled upstream via yara_matches table
    if not decompile and not force and not yara_scan and is_in_db(sha256, index_prefix):
        logger.info(f"File already in database: {sha256}, {filepath}")
        return ImportResult.SKIPPED, filetype

    logger.debug(f"Analysing: {filepath}")
    logger.info(f"Filetype: {filetype}")
    logger.info(f"Filesize: {len(data)/1024:.2f} Kilobytes")

    results = {result: 0 for result in ImportResult}
    is_packed = _is_packed(sha256, index_prefix, logger)

    # Normalize selected_modules once. CLI default is the string "all"; some
    # callers pass a comma-separated list; programmatic callers may pass None.
    # All downstream gates expect a list, so convert here and only here. The
    # historical normalization buried at the top of the non-decompile branch
    # only saw it for one code path, leaving the decompile branch's IOC gates
    # operating on a raw string (substring matching, fragile).
    if isinstance(selected_modules, str):
        selected_modules = [m.strip() for m in selected_modules.split(",")]
    if not selected_modules:
        selected_modules = ['all']

    # Single source of truth for "should the IOC extractor run for this
    # sample". Trips on either 'all' or the explicit class name. Used at
    # every IOC call site (APK / FAT-macho slice / regular binary / JS) so
    # `--modules JSFeaturesExtractor` (or any non-IOC single selection) gets
    # the obvious "skip everything else" semantics globally, not per-format.
    ioc_module_selected = (
        'all' in selected_modules
        or 'IOCExtractorFromResults' in selected_modules
    )

    try:
        # Handle decompilation if needed
        if decompile:
            logger.debug("Running decompilation")
            if filetype == "apk":
                # APK decompilation uses JADX/apktool, not Binary Ninja
                logger.debug("APK file detected, using DecompileAPK")
                if not force and is_in_code_db(sha256, filetype=filetype):
                    logger.info(f"APK {sha256} already disassembled (smali), skipping")
                else:
                    try:
                        DecompileAPK = get_extractor_class("DecompileAPK")
                        decompiler = DecompileAPK(
                            filepath,
                            logger,
                            exporters=exporters,
                            index_prefix=index_prefix,
                            filetype="apk",
                        )
                        result = decompiler.export_data()
                        if result:
                            results[ImportResult.CORRECTLY] += 1

                            # Run IOC extraction from in-memory analysis results.
                            # Gated on the global ioc_module_selected so
                            # `--modules <SomethingElse>` skips IOC scraping
                            # consistently with the JS / binary paths.
                            if decompiler.analysis_results and ioc_module_selected:
                                try:
                                    from redb.extractors.ioc_extractor.ioc_extractor import IOCExtractorFromResults
                                    from redb.extractors.ioc_extractor.standalone_ioc_extractor import IOCType

                                    ioc_extractor = IOCExtractorFromResults(
                                        analysis_results=decompiler.analysis_results,
                                        sha256=sha256,
                                        log=logger,
                                        exporters=exporters,
                                        index_prefix=index_prefix,
                                        suppress_types={IOCType.FQDN},
                                    )
                                    ioc_extractor.export_data()
                                except Exception as e:
                                    logger.warning(f"APK IOC extraction failed (non-fatal): {e}")
                            elif decompiler.analysis_results and not ioc_module_selected:
                                logger.debug(
                                    "Skipping APK IOC extraction (not in selected_modules)"
                                )

                        elif result is False:
                            results[ImportResult.FAILED] += 1
                    except Exception as e:
                        logger.error(f"Error in DecompileAPK: {str(e)}")
                        results[ImportResult.FAILED] += 1
            elif is_packed:
                logger.info(f"File {filepath} is packed, skipping decompilation")
            else:
                from redb.extractors.decompiler import DecompileBinja
                # Check if this is a FAT Mach-O binary - need to decompile each slice separately
                if filetype == "macho":
                    try:
                        import machofile
                        macho = machofile.UniversalMachO(filepath)
                        macho.parse()
                        architectures = macho.get_architectures()
                        is_fat = len(architectures) > 1

                        if is_fat:
                            # FAT Mach-O: extract slices and decompile each one
                            logger.info(f"FAT Mach-O detected with {len(architectures)} architectures, decompiling each slice")

                            with tempfile.TemporaryDirectory() as slice_dir:
                                slices = extract_fat_slices(macho, slice_dir, logger)

                                for arch_name, slice_path, slice_sha256, slice_md5, slice_sha1 in slices:
                                    logger.info(f"Decompiling FAT slice {arch_name} ({slice_sha256})")

                                    # Check if slice already decompiled (skip check when --force)
                                    if not force and is_in_code_db(slice_sha256, filetype=filetype):
                                        logger.info(f"Slice {arch_name} ({slice_sha256}) already disassembled, skipping")
                                        continue

                                    try:
                                        with DecompileBinja(
                                            slice_path,
                                            logger,
                                            exporters=exporters,
                                            index_prefix=index_prefix,
                                            filetype=filetype,
                                            decompile_modules=decompile_modules,
                                        ) as decompiler:
                                            result = decompiler.export_data()
                                            if result:
                                                results[ImportResult.CORRECTLY] += 1

                                                # Run IOC extraction for this slice
                                                # IOC uses decompiled functions + strings, so it runs
                                                # automatically when those modules produce data.
                                                # Two independent gates apply:
                                                #   - decompile_modules (runtime sub-stage selector)
                                                #   - ioc_module_selected (global --modules gate
                                                #     shared across all formats)
                                                run_ioc = (
                                                    not decompile_modules
                                                    or "all" in decompile_modules
                                                    or "decompilation" in decompile_modules
                                                    or "strings" in decompile_modules
                                                ) and ioc_module_selected
                                                if run_ioc:
                                                    try:
                                                        from redb.extractors.ioc_extractor.ioc_extractor import IOCExtractorFromResults

                                                        if decompiler.analysis_results:
                                                            ioc_extractor = IOCExtractorFromResults(
                                                                analysis_results=decompiler.analysis_results,
                                                                sha256=slice_sha256,
                                                                log=logger,
                                                                exporters=exporters,
                                                                index_prefix=index_prefix,
                                                            )
                                                            ioc_extractor.export_data()
                                                        else:
                                                            logger.debug(f"No analysis results available for IOC extraction (slice {arch_name})")
                                                    except Exception as e:
                                                        logger.warning(f"IOC extraction failed for slice {arch_name} (non-fatal): {e}")
                                                else:
                                                    logger.debug(f"Skipping IOC extraction for slice {arch_name} (not relevant for selected decompile modules)")

                                            elif result is False:
                                                results[ImportResult.FAILED] += 1
                                    except Exception as e:
                                        logger.error(f"Error decompiling slice {arch_name}: {str(e)}")
                                        results[ImportResult.FAILED] += 1

                            # Skip the normal decompilation path since we handled FAT
                        else:
                            # Single-arch Mach-O: proceed with normal decompilation below
                            pass
                    except Exception as e:
                        logger.error(f"Error parsing Mach-O for FAT detection: {str(e)}")
                        # Fall through to normal decompilation as fallback

                # Normal decompilation for non-FAT binaries (or single-arch Mach-O)
                if not (filetype == "macho" and 'is_fat' in locals() and is_fat):
                    # Check if already decompiled (applies to all file types, skip when --force)
                    if not force and is_in_code_db(sha256, filetype=filetype):
                        logger.info(f"File {sha256} already disassembled, skipping")
                    else:
                        try:
                            with DecompileBinja(
                                filepath,
                                logger,
                                exporters=exporters,
                                index_prefix=index_prefix,
                                filetype=filetype,
                                decompile_modules=decompile_modules,
                            ) as decompiler:
                                result = decompiler.export_data()
                                if result:
                                    results[ImportResult.CORRECTLY] += 1

                                    # Run IOC extraction from in-memory analysis results.
                                    # IOC uses decompiled functions + strings, so it runs
                                    # automatically when those modules produce data.
                                    # AND-gated with the global --modules selector so
                                    # `--modules <SomethingElse>` skips IOC scraping.
                                    run_ioc = (
                                        not decompile_modules
                                        or "all" in decompile_modules
                                        or "decompilation" in decompile_modules
                                        or "strings" in decompile_modules
                                    ) and ioc_module_selected
                                    if run_ioc:
                                        try:
                                            from redb.extractors.ioc_extractor.ioc_extractor import IOCExtractorFromResults

                                            # Use analysis results directly (no ClickHouse delay needed)
                                            if decompiler.analysis_results:
                                                ioc_extractor = IOCExtractorFromResults(
                                                    analysis_results=decompiler.analysis_results,
                                                    sha256=sha256,
                                                    log=logger,
                                                    exporters=exporters,
                                                    index_prefix=index_prefix,
                                                )
                                                ioc_extractor.export_data()
                                            else:
                                                logger.debug("No analysis results available for IOC extraction")
                                        except Exception as e:
                                            # Non-fatal: log warning but don't fail decompilation
                                            logger.warning(f"IOC extraction failed (non-fatal): {e}")
                                    else:
                                        logger.debug("Skipping IOC extraction (not relevant for selected decompile modules)")

                                elif decompiler.is_dotnet():
                                    # .NET binaries are intentionally skipped, not failed
                                    logger.info("Skipping .NET binary - decompilation not supported")
                                elif result is False:
                                    # Only count as failed if explicitly False, not None
                                    results[ImportResult.FAILED] += 1
                                # result is None means no data to export (not a failure)
                        except Exception as e:
                            logger.error(f"Error in decompilation: {str(e)}")
                            results[ImportResult.FAILED] += 1
        elif yara_scan:
            # Handle YARA scanning only mode
            logger.debug("Running YARA scan only")
            try:
                extractor = YaraExtractor(
                    filepath,
                    logger,
                    exporters=exporters,
                    index_prefix=index_prefix
                )
                result = extractor.export_data()
                # YARA scan is successful even if no rules matched
                # result being None/False just means no matches, not a failure
                results[ImportResult.CORRECTLY] += 1
                if result:
                    logger.info(f"YARA scan completed with matches")
                else:
                    logger.info(f"YARA scan completed with no matches")
            except Exception as e:
                logger.error(f"Error in YARA scan: {str(e)}")
                results[ImportResult.FAILED] += 1
        else:
            # selected_modules was already normalized at the top of
            # process_binary_file — no per-branch normalization needed.

            # 1. First run DIE if selected or 'all'
            # Skip for Mach-O - DIE runs per-slice in the Mach-O handling section
            # Skip for JavaScript - DIE does not support text-based formats
            if filetype not in ("macho", "javascript"):
                if 'all' in selected_modules or 'DIEExtractor' in selected_modules:
                    try:
                        logger.debug("Running DIEExtractor")
                        extractor = DIEExtractor(
                            filepath,
                            logger,
                            exporters=exporters,
                            index_prefix=index_prefix
                        )
                        result = extractor.export_data()
                        is_packed = bool(result) if result is not None else None
                        if result:
                            results[ImportResult.CORRECTLY] += 1
                        elif result is False:
                            results[ImportResult.FAILED] += 1
                    except Exception as e:
                        logger.error(f"Error in DIEExtractor: {str(e)}")
                        results[ImportResult.FAILED] += 1

            # 2. Then run BasicProperties if selected or 'all'
            # Skip for Mach-O - BasicProperties runs for FAT container + each slice in the Mach-O handling section
            # Skip for APK - BasicProperties runs in the APK handling section
            if filetype not in ("macho", "apk", "javascript"):
                if 'all' in selected_modules or 'BasicPropertiesExtractor' in selected_modules:
                    try:
                        logger.debug("Running BasicPropertiesExtractor")
                        extractor = BasicPropertiesExtractor(
                            filepath,
                            logger,
                            exporters=exporters,
                            index_prefix=index_prefix,
                            first_seen=first_seen,
                        )
                        if is_packed is not None:
                            extractor.is_packed = is_packed
                        else:
                            extractor.is_packed = _is_packed(sha256, index_prefix, logger)
                        result = extractor.export_data()
                        if result:
                            results[ImportResult.CORRECTLY] += 1
                        elif result is False:
                            results[ImportResult.FAILED] += 1
                    except Exception as e:
                        logger.error(f"Error in BasicPropertiesExtractor: {str(e)}")
                        results[ImportResult.FAILED] += 1

            # 3. Run HashExtractor if selected or 'all'
            # Skip for Mach-O - HashExtractor runs per-slice in the Mach-O handling section
            # Skip for APK - HashExtractor runs in the APK handling section
            if filetype not in ("macho", "apk", "javascript"):
                if 'all' in selected_modules or 'HashExtractor' in selected_modules:
                    try:
                        logger.debug("Running HashExtractor")
                        extractor = HashExtractor(
                            filepath,
                            logger,
                            exporters=exporters,
                            index_prefix=index_prefix
                        )
                        result = extractor.export_data()
                        if result:
                            results[ImportResult.CORRECTLY] += 1
                        elif result is False:
                            results[ImportResult.FAILED] += 1
                    except Exception as e:
                        logger.error(f"Error in HashExtractor: {str(e)}")
                        results[ImportResult.FAILED] += 1

            # 4. Run CAPA only for supported formats (PE, ELF, .NET)
            if filetype in ('pebin', 'elf'):
                if 'all' in selected_modules or 'CAPAExtractor' in selected_modules:
                    try:
                        logger.debug("Running CAPAExtractor")
                        extractor = CAPAExtractor(
                            filepath,
                            logger,
                            exporters=exporters,
                            index_prefix=index_prefix
                        )
                        result = extractor.export_data()
                        if result:
                            results[ImportResult.CORRECTLY] += 1
                        elif result is False:
                            results[ImportResult.FAILED] += 1
                    except Exception as e:
                        logger.error(f"Error in CAPAExtractor: {str(e)}")
                        results[ImportResult.FAILED] += 1

            # 5. Run Malcontent for supported formats (configurable via MALCONTENT_FORMATS)
            malcontent_formats = os.getenv("MALCONTENT_FORMATS", "pebin,elf,macho,apk,javascript").split(",")
            if filetype in malcontent_formats:
                if 'all' in selected_modules or 'MalcontentExtractor' in selected_modules:
                    try:
                        logger.debug("Running MalcontentExtractor")
                        extractor = MalcontentExtractor(
                            filepath,
                            logger,
                            exporters=exporters,
                            index_prefix=index_prefix
                        )
                        result = extractor.export_data()
                        if result:
                            results[ImportResult.CORRECTLY] += 1
                        elif result is False:
                            results[ImportResult.FAILED] += 1
                    except Exception as e:
                        logger.error(f"Error in MalcontentExtractor: {str(e)}")
                        results[ImportResult.FAILED] += 1

            # 6. Handle filetype-specific modules
            if filetype == "pebin":
                pe_modules = [m for m in get_filetype_modules("pebin")
                              if m.__name__ != "PEDotNetExtractor"]

                for module in pe_modules:
                    if 'all' in selected_modules or module.__name__ in selected_modules:
                        try:
                            logger.debug(f"Running {module.__name__}")
                            extractor = module(
                                filepath,
                                logger,
                                exporters=exporters,
                                index_prefix=index_prefix
                            )
                            result = extractor.export_data()
                            if result:
                                results[ImportResult.CORRECTLY] += 1
                            elif result is False:
                                results[ImportResult.FAILED] += 1
                            # result is None means no data to export (not a failure)
                        except Exception as e:
                            logger.error(f"Error in {module.__name__}: {str(e)}")
                            results[ImportResult.FAILED] += 1

                # Handle PEDotNetExtractor separately due to its special check
                if 'all' in selected_modules or 'PEDotNetExtractor' in selected_modules:
                    if check_dotnet(filepath, logger):
                        try:
                            logger.debug("Running PEDotNetExtractor")
                            PEDotNetExtractor = get_extractor_class("PEDotNetExtractor")
                            extractor = PEDotNetExtractor(
                                filepath,
                                logger,
                                exporters=exporters,
                                index_prefix=index_prefix
                            )
                            result = extractor.export_data()
                            if result:
                                results[ImportResult.CORRECTLY] += 1
                            elif result is False:
                                results[ImportResult.FAILED] += 1
                            # result is None means no data to export (not a failure)
                        except Exception as e:
                            logger.error(f"Error in PEDotNetExtractor: {str(e)}")
                            results[ImportResult.FAILED] += 1
            elif filetype == "elf":
                logger.debug("ELF file detected")
                elf_modules = get_filetype_modules("elf")

                # Parse ELF file once and share across all extractors for efficiency
                try:
                    from elftools.elf.elffile import ELFFile
                    with open(filepath, 'rb') as f:
                        elf_file = ELFFile(f)
                        if not elf_file:
                            logger.error(f"Failed to parse ELF file: {filepath}")
                        else:
                            logger.debug("ELF file parsed successfully, sharing across extractors")

                            for module in elf_modules:
                                if 'all' in selected_modules or module.__name__ in selected_modules:
                                    try:
                                        logger.debug(f"Running {module.__name__}")
                                        # Pass the pre-parsed ELF object to avoid re-parsing
                                        extractor = module(
                                            filepath,
                                            logger,
                                            exporters=exporters,
                                            index_prefix=index_prefix,
                                            elf=elf_file
                                        )
                                        result = extractor.export_data()
                                        if result:
                                            results[ImportResult.CORRECTLY] += 1
                                        elif result is False:
                                            results[ImportResult.FAILED] += 1
                                        # result is None means no data to export (not a failure)
                                    except Exception as e:
                                        logger.error(f"Error in {module.__name__}: {str(e)}")
                                        results[ImportResult.FAILED] += 1
                except Exception as e:
                    logger.error(f"Failed to parse ELF file {filepath}: {str(e)}")
                    # Fallback to individual parsing if shared parsing fails
                    for module in elf_modules:
                        if 'all' in selected_modules or module.__name__ in selected_modules:
                            try:
                                logger.debug(f"Running {module.__name__} (fallback mode)")
                                extractor = module(
                                    filepath,
                                    logger,
                                    exporters=exporters,
                                    index_prefix=index_prefix
                                )
                                result = extractor.export_data()
                                if result:
                                    results[ImportResult.CORRECTLY] += 1
                                elif result is False:
                                    results[ImportResult.FAILED] += 1
                                # result is None means no data to export (not a failure)
                            except Exception as e:
                                logger.error(f"Error in {module.__name__}: {str(e)}")
                                results[ImportResult.FAILED] += 1

            elif filetype == "macho":
                logger.debug("MACHO file detected")
                macho_modules = get_filetype_modules("macho")

                # Parse MachO file once and share across all extractors for efficiency
                try:
                    import machofile
                    macho = machofile.UniversalMachO(filepath)
                    macho.parse()
                    logger.debug("MachO file parsed successfully, sharing across extractors")

                    architectures = macho.get_architectures()
                    is_fat = len(architectures) > 1
                    logger.debug(f"MachO architectures: {architectures}, is_fat: {is_fat}")

                    if is_fat:
                        # === FAT BINARY HANDLING ===
                        # Get FAT container hash info
                        fat_general_info = macho.get_general_info()
                        fat_info = fat_general_info.get('fat', {})
                        fat_sha256 = fat_info.get('SHA256')
                        fat_md5 = fat_info.get('MD5')
                        fat_sha1 = fat_info.get('SHA1')
                        logger.info(f"Processing FAT binary with {len(architectures)} architectures: {fat_sha256}")

                        # Collect slice info: SHA256, architecture, and filetype for each slice
                        child_sha256_list = []
                        child_architecture_list = []
                        child_filetype_list = []
                        for arch_name in architectures:
                            arch_info = macho.get_general_info(arch=arch_name)
                            if arch_info and arch_info.get('SHA256'):
                                child_sha256_list.append(arch_info.get('SHA256'))
                                child_architecture_list.append(arch_name)
                                # For FAT Mach-O, all slices are macho type
                                child_filetype_list.append('macho')
                        logger.debug(f"FAT container children: sha256={child_sha256_list}, arch={child_architecture_list}")

                        # 1. Insert FAT container into basic_properties (parent_sha256=NULL, is_fat=True)
                        if 'all' in selected_modules or 'BasicPropertiesExtractor' in selected_modules:
                            try:
                                logger.debug("Running BasicPropertiesExtractor for FAT container")
                                fat_extractor = BasicPropertiesExtractor(
                                    filepath,
                                    logger,
                                    exporters=exporters,
                                    index_prefix=index_prefix,
                                    parent_sha256=None,
                                    precomputed_hashes={'SHA256': fat_sha256, 'MD5': fat_md5, 'SHA1': fat_sha1},
                                    is_fat=True,
                                    child_sha256=child_sha256_list,
                                    child_architecture=child_architecture_list,
                                    child_filetype=child_filetype_list,
                                    first_seen=first_seen,
                                )
                                if is_packed is not None:
                                    fat_extractor.is_packed = is_packed
                                result = fat_extractor.export_data()
                                if result:
                                    results[ImportResult.CORRECTLY] += 1
                                elif result is False:
                                    results[ImportResult.FAILED] += 1
                            except Exception as e:
                                logger.error(f"Error in BasicPropertiesExtractor (FAT container): {str(e)}")
                                results[ImportResult.FAILED] += 1

                        # 1b. HashExtractor for FAT container (with macho object for similarity hashes)
                        if 'all' in selected_modules or 'HashExtractor' in selected_modules:
                            try:
                                logger.debug("Running HashExtractor for FAT container")
                                fat_hash_extractor = HashExtractor(
                                    filepath,
                                    logger,
                                    exporters=exporters,
                                    index_prefix=index_prefix,
                                    macho=macho
                                )
                                result = fat_hash_extractor.export_data()
                                if result:
                                    results[ImportResult.CORRECTLY] += 1
                                elif result is False:
                                    results[ImportResult.FAILED] += 1
                            except Exception as e:
                                logger.error(f"Error in HashExtractor (FAT container): {str(e)}")
                                results[ImportResult.FAILED] += 1

                        # 2. Extract slices and process each
                        with tempfile.TemporaryDirectory() as slice_dir:
                            slices = extract_fat_slices(macho, slice_dir, logger)

                            for arch_name, slice_path, slice_sha256, slice_md5, slice_sha1 in slices:
                                # 2a. BasicProperties for slice (parent_sha256=fat_sha256)
                                if 'all' in selected_modules or 'BasicPropertiesExtractor' in selected_modules:
                                    try:
                                        logger.debug(f"Running BasicPropertiesExtractor for slice {arch_name}")
                                        slice_extractor = BasicPropertiesExtractor(
                                            slice_path,
                                            logger,
                                            exporters=exporters,
                                            index_prefix=index_prefix,
                                            parent_sha256=fat_sha256,
                                            precomputed_hashes={'SHA256': slice_sha256, 'MD5': slice_md5, 'SHA1': slice_sha1},
                                            first_seen=first_seen,
                                        )
                                        if is_packed is not None:
                                            slice_extractor.is_packed = is_packed
                                        result = slice_extractor.export_data()
                                        if result:
                                            results[ImportResult.CORRECTLY] += 1
                                        elif result is False:
                                            results[ImportResult.FAILED] += 1
                                    except Exception as e:
                                        logger.error(f"Error in BasicPropertiesExtractor (slice {arch_name}): {str(e)}")
                                        results[ImportResult.FAILED] += 1

                                # 2b. DIE for slice
                                if 'all' in selected_modules or 'DIEExtractor' in selected_modules:
                                    try:
                                        logger.debug(f"Running DIEExtractor for slice {arch_name}")
                                        die_extractor = DIEExtractor(
                                            slice_path,
                                            logger,
                                            exporters=exporters,
                                            index_prefix=index_prefix,
                                            precomputed_hashes={'SHA256': slice_sha256, 'MD5': slice_md5, 'SHA1': slice_sha1}
                                        )
                                        result = die_extractor.export_data()
                                        if result:
                                            results[ImportResult.CORRECTLY] += 1
                                        elif result is False:
                                            results[ImportResult.FAILED] += 1
                                    except Exception as e:
                                        logger.error(f"Error in DIEExtractor (slice {arch_name}): {str(e)}")
                                        results[ImportResult.FAILED] += 1

                                # 2c. HashExtractor for slice
                                if 'all' in selected_modules or 'HashExtractor' in selected_modules:
                                    try:
                                        logger.debug(f"Running HashExtractor for slice {arch_name}")
                                        hash_extractor = HashExtractor(
                                            slice_path,
                                            logger,
                                            exporters=exporters,
                                            index_prefix=index_prefix
                                        )
                                        result = hash_extractor.export_data()
                                        if result:
                                            results[ImportResult.CORRECTLY] += 1
                                        elif result is False:
                                            results[ImportResult.FAILED] += 1
                                    except Exception as e:
                                        logger.error(f"Error in HashExtractor (slice {arch_name}): {str(e)}")
                                        results[ImportResult.FAILED] += 1

                        # 3. MachO extractors (already work per-slice via machofile API)
                        for module in macho_modules:
                            if 'all' in selected_modules or module.__name__ in selected_modules:
                                try:
                                    logger.debug(f"Running {module.__name__}")
                                    extractor = module(
                                        filepath,
                                        logger,
                                        exporters=exporters,
                                        index_prefix=index_prefix,
                                        macho=macho
                                    )
                                    result = extractor.export_data()
                                    if result:
                                        results[ImportResult.CORRECTLY] += 1
                                    elif result is False:
                                        results[ImportResult.FAILED] += 1
                                except Exception as e:
                                    logger.error(f"Error in {module.__name__}: {str(e)}")
                                    results[ImportResult.FAILED] += 1

                    else:
                        # === SINGLE-ARCH HANDLING ===
                        # Run DIE for single-arch Mach-O
                        if 'all' in selected_modules or 'DIEExtractor' in selected_modules:
                            try:
                                logger.debug("Running DIEExtractor for single-arch Mach-O")
                                extractor = DIEExtractor(
                                    filepath,
                                    logger,
                                    exporters=exporters,
                                    index_prefix=index_prefix
                                )
                                result = extractor.export_data()
                                is_packed = bool(result) if result is not None else None
                                if result:
                                    results[ImportResult.CORRECTLY] += 1
                                elif result is False:
                                    results[ImportResult.FAILED] += 1
                            except Exception as e:
                                logger.error(f"Error in DIEExtractor: {str(e)}")
                                results[ImportResult.FAILED] += 1

                        # Run BasicProperties for single-arch Mach-O
                        if 'all' in selected_modules or 'BasicPropertiesExtractor' in selected_modules:
                            try:
                                logger.debug("Running BasicPropertiesExtractor for single-arch Mach-O")
                                extractor = BasicPropertiesExtractor(
                                    filepath,
                                    logger,
                                    exporters=exporters,
                                    index_prefix=index_prefix,
                                    parent_sha256=None,
                                    first_seen=first_seen,
                                )
                                if is_packed is not None:
                                    extractor.is_packed = is_packed
                                result = extractor.export_data()
                                if result:
                                    results[ImportResult.CORRECTLY] += 1
                                elif result is False:
                                    results[ImportResult.FAILED] += 1
                            except Exception as e:
                                logger.error(f"Error in BasicPropertiesExtractor: {str(e)}")
                                results[ImportResult.FAILED] += 1

                        # Run HashExtractor for single-arch Mach-O (with macho object for similarity hashes)
                        if 'all' in selected_modules or 'HashExtractor' in selected_modules:
                            try:
                                logger.debug("Running HashExtractor for single-arch Mach-O")
                                extractor = HashExtractor(
                                    filepath,
                                    logger,
                                    exporters=exporters,
                                    index_prefix=index_prefix,
                                    macho=macho
                                )
                                result = extractor.export_data()
                                if result:
                                    results[ImportResult.CORRECTLY] += 1
                                elif result is False:
                                    results[ImportResult.FAILED] += 1
                            except Exception as e:
                                logger.error(f"Error in HashExtractor: {str(e)}")
                                results[ImportResult.FAILED] += 1

                        # Run MachO-specific extractors
                        for module in macho_modules:
                            if 'all' in selected_modules or module.__name__ in selected_modules:
                                try:
                                    logger.debug(f"Running {module.__name__}")
                                    extractor = module(
                                        filepath,
                                        logger,
                                        exporters=exporters,
                                        index_prefix=index_prefix,
                                        macho=macho
                                    )
                                    result = extractor.export_data()
                                    if result:
                                        results[ImportResult.CORRECTLY] += 1
                                    elif result is False:
                                        results[ImportResult.FAILED] += 1
                                except Exception as e:
                                    logger.error(f"Error in {module.__name__}: {str(e)}")
                                    results[ImportResult.FAILED] += 1

                except Exception as e:
                    logger.error(f"Failed to parse MachO file {filepath}: {str(e)}")

                    # Run generic extractors — these don't need machofile
                    # and must succeed even for corrupted/unparseable Mach-O files
                    if 'all' in selected_modules or 'BasicPropertiesExtractor' in selected_modules:
                        try:
                            logger.debug("Running BasicPropertiesExtractor for Mach-O (fallback)")
                            extractor = BasicPropertiesExtractor(
                                filepath,
                                logger,
                                exporters=exporters,
                                index_prefix=index_prefix,
                                parent_sha256=None,
                                first_seen=first_seen,
                            )
                            if is_packed is not None:
                                extractor.is_packed = is_packed
                            result = extractor.export_data()
                            if result:
                                results[ImportResult.CORRECTLY] += 1
                            elif result is False:
                                results[ImportResult.FAILED] += 1
                        except Exception as e:
                            logger.error(f"Error in BasicPropertiesExtractor (Mach-O fallback): {str(e)}")
                            results[ImportResult.FAILED] += 1

                    if 'all' in selected_modules or 'HashExtractor' in selected_modules:
                        try:
                            logger.debug("Running HashExtractor for Mach-O (fallback)")
                            extractor = HashExtractor(
                                filepath,
                                logger,
                                exporters=exporters,
                                index_prefix=index_prefix,
                            )
                            result = extractor.export_data()
                            if result:
                                results[ImportResult.CORRECTLY] += 1
                            elif result is False:
                                results[ImportResult.FAILED] += 1
                        except Exception as e:
                            logger.error(f"Error in HashExtractor (Mach-O fallback): {str(e)}")
                            results[ImportResult.FAILED] += 1

                    # Fallback to individual parsing for format-specific extractors
                    for module in macho_modules:
                        if 'all' in selected_modules or module.__name__ in selected_modules:
                            try:
                                logger.debug(f"Running {module.__name__} (fallback mode)")
                                extractor = module(
                                    filepath,
                                    logger,
                                    exporters=exporters,
                                    index_prefix=index_prefix
                                )
                                result = extractor.export_data()
                                if result:
                                    results[ImportResult.CORRECTLY] += 1
                                elif result is False:
                                    results[ImportResult.FAILED] += 1
                            except Exception as e:
                                logger.error(f"Error in {module.__name__}: {str(e)}")
                                results[ImportResult.FAILED] += 1

            elif filetype == "apk":
                logger.debug("APK file detected")
                apk_modules = get_filetype_modules("apk")

                # Run generic extractors first — these don't need androguard
                # and must succeed even for corrupted/invalid APKs
                if 'all' in selected_modules or 'BasicPropertiesExtractor' in selected_modules:
                    try:
                        logger.debug("Running BasicPropertiesExtractor for APK")
                        extractor = BasicPropertiesExtractor(
                            filepath,
                            logger,
                            exporters=exporters,
                            index_prefix=index_prefix,
                            parent_sha256=None,
                            first_seen=first_seen,
                        )
                        if is_packed is not None:
                            extractor.is_packed = is_packed
                        result = extractor.export_data()
                        if result:
                            results[ImportResult.CORRECTLY] += 1
                        elif result is False:
                            results[ImportResult.FAILED] += 1
                    except Exception as e:
                        logger.error(f"Error in BasicPropertiesExtractor (APK): {str(e)}")
                        results[ImportResult.FAILED] += 1

                if 'all' in selected_modules or 'HashExtractor' in selected_modules:
                    try:
                        logger.debug("Running HashExtractor for APK")
                        extractor = HashExtractor(
                            filepath,
                            logger,
                            exporters=exporters,
                            index_prefix=index_prefix,
                        )
                        result = extractor.export_data()
                        if result:
                            results[ImportResult.CORRECTLY] += 1
                        elif result is False:
                            results[ImportResult.FAILED] += 1
                    except Exception as e:
                        logger.error(f"Error in HashExtractor (APK): {str(e)}")
                        results[ImportResult.FAILED] += 1

                # Parse APK file once and share across all APK-specific extractors
                try:
                    from androguard.core.apk import APK
                    apk = APK(filepath)
                    logger.debug("APK file parsed successfully, sharing across extractors")

                    # Run APK-specific extractors
                    for module in apk_modules:
                        if 'all' in selected_modules or module.__name__ in selected_modules:
                            try:
                                logger.debug(f"Running {module.__name__}")
                                extractor = module(
                                    filepath,
                                    logger,
                                    exporters=exporters,
                                    index_prefix=index_prefix,
                                    apk=apk
                                )
                                result = extractor.export_data()
                                if result:
                                    results[ImportResult.CORRECTLY] += 1
                                elif result is False:
                                    results[ImportResult.FAILED] += 1
                            except Exception as e:
                                logger.error(f"Error in {module.__name__}: {str(e)}")
                                results[ImportResult.FAILED] += 1

                except Exception as e:
                    logger.error(f"Failed to parse APK file {filepath}: {str(e)}")
                    # Fallback to individual parsing if shared parsing fails
                    for module in apk_modules:
                        if 'all' in selected_modules or module.__name__ in selected_modules:
                            try:
                                logger.debug(f"Running {module.__name__} (fallback mode)")
                                extractor = module(
                                    filepath,
                                    logger,
                                    exporters=exporters,
                                    index_prefix=index_prefix
                                )
                                result = extractor.export_data()
                                if result:
                                    results[ImportResult.CORRECTLY] += 1
                                elif result is False:
                                    results[ImportResult.FAILED] += 1
                            except Exception as e:
                                logger.error(f"Error in {module.__name__}: {str(e)}")
                                results[ImportResult.FAILED] += 1
            elif filetype == "javascript":
                logger.debug("JavaScript file detected")
                js_modules = get_filetype_modules("javascript")

                # Run generic extractors first
                if 'all' in selected_modules or 'BasicPropertiesExtractor' in selected_modules:
                    try:
                        logger.debug("Running BasicPropertiesExtractor for JavaScript")
                        extractor = BasicPropertiesExtractor(
                            filepath,
                            logger,
                            exporters=exporters,
                            index_prefix=index_prefix,
                            first_seen=first_seen,
                            parent_sha256=None,
                        )
                        result = extractor.export_data()
                        if result:
                            results[ImportResult.CORRECTLY] += 1
                        elif result is False:
                            results[ImportResult.FAILED] += 1
                    except Exception as e:
                        logger.error(f"Error in BasicPropertiesExtractor (JS): {str(e)}")
                        results[ImportResult.FAILED] += 1

                if 'all' in selected_modules or 'HashExtractor' in selected_modules:
                    try:
                        logger.debug("Running HashExtractor for JavaScript")
                        extractor = HashExtractor(
                            filepath,
                            logger,
                            exporters=exporters,
                            index_prefix=index_prefix,
                        )
                        result = extractor.export_data()
                        if result:
                            results[ImportResult.CORRECTLY] += 1
                        elif result is False:
                            results[ImportResult.FAILED] += 1
                    except Exception as e:
                        logger.error(f"Error in HashExtractor (JS): {str(e)}")
                        results[ImportResult.FAILED] += 1

                # Build one JSContext per sample and thread it through every
                # JS extractor: disk read + decode + scan + AST happen once.
                try:
                    from redb.extractors.js_extractors.js_context import JSContext

                    js_ctx = JSContext.from_path(
                        filepath, log=logger, content_type=filetype
                    )
                    source = js_ctx.source
                    logger.debug("JS source loaded, sharing across extractors")

                    for module in js_modules:
                        if 'all' in selected_modules or module.__name__ in selected_modules:
                            try:
                                logger.debug(f"Running {module.__name__}")
                                extractor = module(
                                    filepath,
                                    logger,
                                    exporters=exporters,
                                    index_prefix=index_prefix,
                                    context=js_ctx,
                                )
                                result = extractor.export_data()
                                if result:
                                    results[ImportResult.CORRECTLY] += 1
                                elif result is False:
                                    results[ImportResult.FAILED] += 1
                            except Exception as e:
                                logger.error(f"Error in {module.__name__}: {str(e)}")
                                results[ImportResult.FAILED] += 1

                    # Run IOC extraction using the shared IOC extractor. Gated
                    # on the global `ioc_module_selected` flag computed at the
                    # top of process_binary_file so `--modules
                    # JSFeaturesExtractor` (or any non-IOC single selection)
                    # skips IOC scraping for every format consistently.
                    if ioc_module_selected:
                        try:
                            import hashlib as _hashlib
                            from redb.extractors.ioc_extractor.ioc_extractor import IOCExtractorFromResults

                            # Three IOC surfaces are fed in for JS samples:
                            #   text_raw          — the source as-is on disk
                            #   text_normalized   — the deobfuscated/beautified
                            #                       form (only when it differs
                            #                       from raw; avoids double-
                            #                       scraping byte-identical
                            #                       jsbeautifier output)
                            #   strings           — the per-string decodings
                            #                       from JSStringsExtractor
                            #                       (hex/unicode/charcode/
                            #                       base64/concat unpacked into
                            #                       plaintext; empty when
                            #                       JSStringsExtractor was
                            #                       excluded via --modules)
                            text_raw_entries = [{
                                "content": source,
                                "content_hash": sha256,
                            }]
                            text_normalized_entries = []
                            deobf_text, _ = js_ctx.deobfuscated
                            if deobf_text and deobf_text != source:
                                text_normalized_entries.append({
                                    "content": deobf_text,
                                    "content_hash": _hashlib.sha256(
                                        deobf_text.encode("utf-8")
                                    ).hexdigest(),
                                })

                            decoded_strings = js_ctx.decoded_strings or []
                            strings_entries = [
                                {
                                    "string": s.get("string", ""),
                                    "string_offset": s.get("string_offset", 0),
                                }
                                for s in decoded_strings
                            ]

                            analysis_results = {
                                "strings": strings_entries,
                                "text_raw": text_raw_entries,
                                "text_normalized": text_normalized_entries,
                            }
                            ioc_extractor = IOCExtractorFromResults(
                                analysis_results=analysis_results,
                                sha256=sha256,
                                log=logger,
                                exporters=exporters,
                                index_prefix=index_prefix,
                                # JS-context FQDN filter: rejects candidates
                                # that are JS object-access syntax
                                # (this.foo.bar, process.id, lib.so,
                                # Component.name, ...) which otherwise drown
                                # out real C2 hostnames. See JS_FP_TLDS /
                                # JS_FP_SLDS in standalone_ioc_extractor.
                                js_context=True,
                            )
                            ioc_extractor.export_data()
                        except Exception as e:
                            logger.warning(f"JS IOC extraction failed (non-fatal): {e}")
                    else:
                        logger.debug(
                            "Skipping JS IOC extraction (not in selected_modules)"
                        )

                except Exception as e:
                    logger.error(f"Failed to read JS file {filepath}: {str(e)}")
                    results[ImportResult.FAILED] += 1

            else:
                logger.info(f"Unsupported Filetype: {filetype}")
                return ImportResult.SKIPPED, filetype

            # Handle YARA scanning when with_yara is enabled (combined features + YARA mode)
            if with_yara:
                logger.debug("Running YARA scan (with_yara mode)")
                try:
                    extractor = YaraExtractor(
                        filepath,
                        logger,
                        exporters=exporters,
                        index_prefix=index_prefix
                    )
                    result = extractor.export_data()
                    # YARA scan is successful even if no rules matched
                    results[ImportResult.CORRECTLY] += 1
                    if result:
                        logger.info(f"YARA scan completed with matches")
                    else:
                        logger.info(f"YARA scan completed with no matches")
                except Exception as e:
                    logger.error(f"Error in YARA scan (with_yara): {str(e)}")
                    results[ImportResult.FAILED] += 1

        elapsed_time = time.time() - start_time
        logger.info(f"Total processing time: {elapsed_time:.2f} seconds")

        # Determine final status
        if all(v == 0 for v in results.values()):
            return ImportResult.SKIPPED, filetype
        elif results[ImportResult.CORRECTLY] > 0 and results[ImportResult.FAILED] == 0:
            return ImportResult.CORRECTLY, filetype
        elif results[ImportResult.FAILED] > 0 and results[ImportResult.CORRECTLY] == 0:
            return ImportResult.FAILED, filetype
        elif results[ImportResult.CORRECTLY] > 0 and results[ImportResult.FAILED] > 0:
            # Some succeeded, some failed - this is PARTIALLY
            return ImportResult.PARTIALLY, filetype
        elif results[ImportResult.CORRECTLY] > 0:
            # Only successes, no failures
            return ImportResult.CORRECTLY, filetype
        else:
            # Only failures, no successes
            return ImportResult.FAILED, filetype
    finally:
        # Close client after all modules are done
        if clickhouse_client:
            try:
                clickhouse_client.close()
            except:
                pass


def worker(args):
    """
    Worker function for local files with direct logging (similar to S3 workers)
    """
    (
        filepath,
        decompile,
        index_prefix,
        log_file,
        file_number,
        total_files,
        selected_modules,
        dry_run,
        yara_scan,
        with_yara,
        force,
        decompile_modules,
    ) = args

    # Setup timeout management - use different timeouts for decompilation
    if decompile:
        REDB_TIMEOUT = int(os.getenv("DECOMPILE_WORKER_TIMEOUT", "2700"))  # 45 minutes for decompilation
    else:
        REDB_TIMEOUT = int(os.getenv("REDB_TIMEOUT", "600"))  # 10 minutes for normal analysis

    # Define timeout handler
    def timeout_handler(signum, frame):
        # Setup a logger
        filename = os.path.basename(filepath)
        worker_pid = os.getpid()
        logger = setup_direct_logger(log_file, filename, worker_pid)

        # Log timeout information
        logger.error(
            f"TIMEOUT: Worker processing {filepath} exceeded {REDB_TIMEOUT} seconds. "
            f"Decompilation mode: {decompile}. Terminating worker."
        )

        # Force clean up before exit to help release resources
        gc.collect()

        # Exit this worker process
        os._exit(1)

    # Register timeout handler
    signal.signal(signal.SIGALRM, timeout_handler)
    signal.alarm(REDB_TIMEOUT)  # Set alarm for REDB_TIMEOUT seconds

    try:
        # Process the file
        result = process_file(
            filepath,
            decompile,
            index_prefix,
            log_file,
            file_number,
            total_files,
            selected_modules,
            dry_run,
            yara_scan,
            with_yara,
            force,
            decompile_modules=decompile_modules,
        )

        # Cancel the alarm
        signal.alarm(0)

        # Force garbage collection before exiting
        gc.collect(2)

        return result

    except Exception as e:
        # Log any exceptions
        filename = os.path.basename(filepath)
        worker_pid = os.getpid()
        logger = setup_direct_logger(log_file, filename, worker_pid)

        logger.error(f"Worker error processing {filepath}: {str(e)}")

        # Force garbage collection before exiting
        gc.collect()

        # Cancel the alarm
        signal.alarm(0)

        return ImportResult.FAILED, None
    finally:
        # Make sure alarm is cancelled
        signal.alarm(0)

        # Final cleanup operations
        gc.collect()

        # Kill any potentially lingering child processes
        try:
            current_proc = psutil.Process()
            for child in current_proc.children(recursive=True):
                try:
                    child.terminate()
                except:
                    pass
        except:
            pass


def direct_s3_worker(s3_bucket, s3_key, temp_dir, decompile, index_prefix, log_file, file_number, total_files, selected_modules, result_queue, dry_run=False, yara_scan=False, with_yara=False, force=False, decompile_modules=None, first_seen=None):
    """
    Worker function that directly puts results in a queue rather than using futures.

    This is the current active S3 worker used in production:
    - Downloads S3 file to local temp directory
    - Processes the file using extractors
    - Puts results directly into a multiprocessing.Queue
    - Uses dynamic timeouts (45min for decompile, 20min for analysis)
    - Designed for multiprocessing.Process() with queue-based result collection
    - Provides better error handling and timeout recovery
    - Used by _process_files_streaming() method for S3 file processing
    """
    import os
    import gc
    import signal
    import time

    # Get our own PID
    worker_pid = os.getpid()

    # Setup direct logging
    logger = setup_direct_logger(log_file, os.path.basename(s3_key), worker_pid)

    # Setup timeout management based on decompile flag
    if decompile:
        REDB_TIMEOUT = int(os.getenv("DECOMPILE_WORKER_TIMEOUT", "2700"))  # 45 minutes for decompilation
    else:
        REDB_TIMEOUT = int(os.getenv("REDB_TIMEOUT", "1200"))  # 20 minutes for normal analysis

    # Define timeout handler
    def timeout_handler(signum, frame):
        logger.error(
            f"TIMEOUT: Worker {worker_pid} processing S3 file {s3_key} exceeded {REDB_TIMEOUT} seconds. "
            f"Decompile mode: {decompile}."
        )
        # Return a failure result
        try:
            result_queue.put({
                "status": "FAILED",
                "filetype": None,
                "worker_pid": worker_pid,
                "s3_key": s3_key,
                "file_number": file_number,
                "error": "Timeout"
            })
        except:
            pass

        # Force exit
        gc.collect()
        os._exit(1)

    # Register timeout handler
    signal.signal(signal.SIGALRM, timeout_handler)
    signal.alarm(REDB_TIMEOUT)

    try:
        logger.info(f"Processing S3 file {file_number}/{total_files}: {s3_key}")

        local_path = download_s3_object(s3_bucket, s3_key, temp_dir, logger)
        if not local_path:
            logger.error(f"Failed to download {s3_key}")
            result_queue.put({
                "status": "FAILED",
                "filetype": None,
                "worker_pid": worker_pid,
                "s3_key": s3_key,
                "file_number": file_number,
                "error": "Download failed"
            })
            return

        # Process the downloaded file
        result, filetype = _process_file_internal(
            local_path, decompile, index_prefix, logger, selected_modules, dry_run, yara_scan, with_yara, force,
            decompile_modules=decompile_modules,
            first_seen=first_seen,
        )

        # Cancel the alarm
        signal.alarm(0)

        # Force garbage collection
        gc.collect()

        # Put result in queue
        result_queue.put({
            "status": result.name if hasattr(result, "name") else str(result),
            "filetype": filetype,
            "worker_pid": worker_pid,
            "s3_key": s3_key,
            "file_number": file_number
        })

    except Exception as e:
        logger.error(f"S3 worker error processing {s3_key}: {str(e)}")

        # Cancel the alarm
        signal.alarm(0)

        # Put error result in queue
        try:
            result_queue.put({
                "status": "FAILED",
                "filetype": None,
                "worker_pid": worker_pid,
                "s3_key": s3_key,
                "file_number": file_number,
                "error": str(e)
            })
        except:
            pass

    finally:
        # Make sure alarm is cancelled
        signal.alarm(0)
        gc.collect()