Radha Thangaraj

34 papers A 1B 3C 1Misc 2Journal 10Unranked 14
YearRankTypeTitle / Venue / Authors
2018 J jnl
Int. J. Syst. Assur. Eng. Manag.
Srikanth Allamsetty, Radha Thangaraj, Thanga Raj Chelliah, Millie Pant
2013 ch.
Handbook of Optimization
Radha Thangaraj, Thanga Raj Chelliah, Millie Pant, Ajith Abraham, Pascal Bouvry
2012 J jnl
Appl. Artif. Intell.
C. Thanga Raj, Radha Thangaraj, Millie Pant, Pascal Bouvry, Ajith Abraham
2012 conf
NaBIC
Radha Thangaraj, Millie Pant, Thanga Raj Chelliah, Ajith Abraham
2012 J jnl
Log. J. IGPL
Radha Thangaraj, Millie Pant, Pascal Bouvry, Ajith Abraham
2011 J jnl
Appl. Artif. Intell.
Radha Thangaraj, Millie Pant, C. Thanga Raj, Atulya K. Nagar
2011 conf
SocProS (1)
S. Anil Chandrakanth, Thanga Raj Chelliah, S. P. Srivastava, Radha Thangaraj
2011 J jnl
Log. J. IGPL
Radha Thangaraj, Thanga Raj Chelliah, Millie Pant, Ajith Abraham, Crina Grosan
2011 J jnl
Appl. Math. Comput.
Radha Thangaraj, Millie Pant, Ajith Abraham, Pascal Bouvry
2010 B conf
SMC
Radha Thangaraj, Millie Pant, Ajith Abraham, Kusum Deep, Václav Snásel
2010 J jnl
Appl. Math. Comput.
Radha Thangaraj, Millie Pant, Ajith Abraham
2010 J jnl
Eng. Appl. Artif. Intell.
Radha Thangaraj, Millie Pant, Kusum Deep
2010 conf
CISIM
Radha Thangaraj, C. Thanga Raj, Pascal Bouvry, Millie Pant, Ajith Abraham
2010 J jnl
Int. J. Artif. Intell. Soft Comput.
Radha Thangaraj, Millie Pant, Atulya K. Nagar, Ved Pal Singh
2010 conf
SEMCCO
Radha Thangaraj, Millie Pant, Pascal Bouvry, Ajith Abraham
2009 conf
NaBIC
Radha Thangaraj, Millie Pant, Ajith Abraham
2009 conf
NaBIC
Radha Thangaraj, Millie Pant, Ajith Abraham
2009 B conf
IEEE Congress on Evolutionary Computation
Millie Pant, Radha Thangaraj, Ajith Abraham, Crina Grosan
2009 conf
NaBIC
Radha Thangaraj, Millie Pant, Ajith Abraham
2009 Misc conf
HAIS
Radha Thangaraj, Millie Pant, Ajith Abraham, Youakim Badr
2009 conf
NaBIC
Radha Thangaraj, Millie Pant, Kusum Deep
2009 J jnl
Fundam. Informaticae
Millie Pant, Radha Thangaraj, Ajith Abraham
2009 ch.
Foundations of Computational Intelligence (3)
Millie Pant, Radha Thangaraj, Ajith Abraham
2008 ch.
Innovations in Hybrid Intelligent Systems
Millie Pant, Radha Thangaraj, Ajith Abraham
2008 A conf
GECCO
Millie Pant, Radha Thangaraj, Ajith Abraham
2008 Misc conf
HAIS
Millie Pant, Radha Thangaraj, Deepti Rani, Ajith Abraham, Dinesh Kumar Srivastava
2008 conf
ICDIM
Millie Pant, Radha Thangaraj, Crina Grosan, Ajith Abraham
2008 B conf
IEEE Congress on Evolutionary Computation
Millie Pant, Radha Thangaraj, Crina Grosan, Ajith Abraham
2008 conf
ISDA (3)
Millie Pant, Radha Thangaraj, Ajith Abraham
2008 conf
Asia International Conference on Modelling and Simulation
Millie Pant, Radha Thangaraj, Ajith Abraham
2008 conf
CISIM
Millie Pant, Radha Thangaraj, Ajith Abraham
2008 conf
DEXA Workshops
Millie Pant, Radha Thangaraj, Ajith Abraham
2008 conf
ICETET
Millie Pant, Radha Thangaraj, Ved Pal Singh, Ajith Abraham
2007 C conf
HIS
Millie Pant, Radha Thangaraj, Ajith Abraham
CLAUDE.md
← Index CLAUDE.md markdown
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## Project Overview

REDB (RationalEdge Samples DB) is a malware analysis framework that extracts features from PE (Portable Executable) files and stores them in ClickHouse database for analysis. It provides a comprehensive set of extractors for analyzing binary samples including PE headers, imports, resources, signatures, and decompiled code.

## Common Commands

### Development Setup
```bash
source venv/bin/activate

# Install dependencies
pip install -r requirements.txt

# Run the main application
python start.py --path /path/to/samples --repo sample_repo --index_prefix redb
```

### Analysis Commands
```bash
# Process a single file
python start.py --path /path/to/binary --repo test --index_prefix redb

# Process from S3 storage
python start.py --s3 --repo malpedia --index_prefix redb

# Process from S3 storage but only a subset of a specific repository
python start.py --s3 --repo "vx-itw" --s3-notes "ITW.0138" --index_prefix redb

# Run only decompilation
python start.py --path /path/to/binary --repo test --index_prefix redb --decompile

# Run specific modules
python start.py --path /path/to/binary --repo test --index_prefix redb --modules "BasicPropertiesExtractor,PEFeaturesExtractor"

# Run as Nomad job (for containerized deployment)
python start.py --nomad-job
```

### Testing
There are no formal unit tests. Testing is done by running the extractors on sample files in the `test_files/` directory.

## Architecture Overview

### Core Components

1. **Ingestor (`redb/ingestor.py`)**: Main orchestrator that handles file processing, multiprocessing, and coordinates extractors
2. **Extractors (`redb/extractors/`)**: Modular analysis components that extract specific features
3. **Database Exporters (`redb/extractors/database_exporters.py`)**: Handle data export to ClickHouse
4. **Settings (`redb/settings/`)**: Configuration management for database connections

### Extractor Architecture

All extractors inherit from the base `Extractor` class and implement:
- `extract()`: Main analysis logic
- `prepare_export_data()`: Format data for database export
- `get_clickhouse_table()`: Return target table name

Available extractors:
- **General**: BasicPropertiesExtractor, HashExtractor, DIEExtractor, CAPAExtractor
- **PE-specific**: PEFeaturesExtractor, PEImportExtractor, PEResourceExtractor, PEOverlayExtractor, PESectionExtractor, PESignatureExtractor, PEDotNetExtractor, PEInconstistencyTestsExtractor, PEExtraFindings
- **ELF**: ELFFeaturesExtractor, ELFSegmentExtractor, ELFSectionExtractor, ELFDependencyExtractor, ELFSymbolExtractor, ELFImportExtractor, ELFExportExtractor, ELFRelocationExtractor, ELFNotesExtractor
- **Mach-O**: MachOFeaturesExtractor, MachOSegmentExtractor, MachOImportExtractor, MachOExportExtractor, MachODylibExtractor, MachOSignatureExtractor
- **APK**: APKFeaturesExtractor, APKManifestExtractor, APKPermissionsExtractor, APKSignatureExtractor, APKDexExtractor, APKResourceExtractor, APKNativeLibExtractor, APKInconsistencyTestsExtractor
- **Decompilation**: DecompileBinja, DecompileAPK

### Database Schema

The project uses a comprehensive ClickHouse schema defined in `redb/redb_schema.yml` with tables for:
- Basic properties (`redb_basic_properties`)
- PE features (`redb_pe_features`, `redb_pe_imports`, `redb_pe_sections`, etc.)
- Decompiled code (`code_binja_decompiled_functions_content`, `code_binja_decompiled_functions_references`)
- CAPA analysis (`redb_capa`, `redb_capa_capabilities`)

Full schema documentation is available in `docs/database_schema.md`.

### Processing Modes

1. **Analysis Mode**: Extracts features using selected modules
2. **Decompile Mode**: Uses Binary Ninja for code decompilation
3. **S3 Mode**: Fetches samples from S3 storage based on catalog queries
4. **Nomad Job Mode**: Processes single jobs using environment variables for containerized deployment

### Configuration

Environment variables are used for configuration:
- Database connection: `CLICKHOUSE_HOST`, `CLICKHOUSE_PORT`, `CLICKHOUSE_USER`, `CLICKHOUSE_PASSWORD`
- S3 storage: `S3_ENDPOINT`, `S3_ACCESS_KEY`, `S3_SECRET_KEY`
- Processing: `BATCH_SIZE`, `REDB_TIMEOUT`, `DECOMPILE_WORKER_TIMEOUT`
- Nomad jobs: `JOB_ID`, `S3_KEY`, `S3_BUCKET`, `WORKER_TYPE`, `CALLBACK_URL`, `ANALYSIS_MODULES`

## Important Implementation Details

### Multiprocessing
- Uses `spawn` method for multiprocessing to avoid memory issues
- Worker processes have timeout handlers to prevent hanging
- Supports both batch processing and streaming processing modes

### Memory Management
- Implements aggressive garbage collection between batches
- Monitors swap usage and restarts worker pools when needed
- Kills stuck processes automatically

### Error Handling
- Comprehensive logging with per-file context
- Graceful handling of corrupted or unsupported files
- Automatic retry logic for database operations

### Security Context
This is a defensive security tool for malware analysis. It processes potentially malicious files in a controlled environment to extract features for detection and analysis purposes.

## Development Notes

- The codebase is optimized for processing large batches of malware samples
- Extractors are designed to be modular and can be run individually or in combination
- Database schema supports both normalized and denormalized views for different query patterns
- S3 integration allows for scalable processing of large malware repositories