Nancy Kanwisher

34 papers A* 2B 4Journal 23Unranked 5
YearRankTypeTitle / Venue / Authors
2025 A* conf
ICLR
Ammar I Marvi, Nancy Kanwisher, Meenakshi Khosla
2025 J jnl
CoRR
Ammar I Marvi, Nancy Kanwisher, Meenakshi Khosla
2025 J jnl
CoRR
Colton Casto, Anna A. Ivanova, Evelina Fedorenko, Nancy Kanwisher
2024 B conf
CogSci
R. T. Pramod, Jessica Chomik, Laura Schulz, Nancy Kanwisher
2024 A* conf
NeurIPS
Tyler Bonnen, Stephanie Fu, Yutong Bai, Thomas P. O'Connell, Yoni Friedman, Nancy Kanwisher, Josh Tenenbaum, Alexei A. Efros
2024 J jnl
CoRR
Tyler Bonnen, Stephanie Fu, Yutong Bai, Thomas P. O'Connell, Yoni Friedman, Nancy Kanwisher, Joshua B. Tenenbaum, Alexei A. Efros
2023 J jnl
CoRR
Thomas P. O'Connell, Tyler Bonnen, Yoni Friedman, Ayush Tewari, Joshua B. Tenenbaum, Vincent Sitzmann, Nancy Kanwisher
2023 J jnl
CoRR
Kyle Mahowald, Anna A. Ivanova, Idan Asher Blank, Nancy Kanwisher, Joshua B. Tenenbaum, Evelina Fedorenko
2022 B conf
CogSci
Heather Kosakowski, Nancy Kanwisher, Rebecca Saxe
2021 conf
NeurIPS Datasets and Benchmarks
Daniel Bear, Elias Wang, Damian Mrowca, Felix J. Binder, Hsiao-Yu Tung, R. T. Pramod, Cameron Holdaway, Sirui Tao, Kevin A. Smith, Fan-Yun Sun, Fei-Fei Li, Nancy Kanwisher, Josh Tenenbaum, Dan Yamins, Judith E. Fan
2021 J jnl
CoRR
Daniel M. Bear, Elias Wang, Damian Mrowca, Felix J. Binder, Hsiau-Yu Fish Tung, R. T. Pramod, Cameron Holdaway, Sirui Tao, Kevin A. Smith, Fan-Yun Sun, Li Fei-Fei, Nancy Kanwisher, Joshua B. Tenenbaum, Daniel L. K. Yamins, Judith E. Fan
2020 B conf
CogSci
Heather Kosakowski, Michael A. Cohen, Nancy Kanwisher, Rebecca Saxe
2020 J jnl
NeuroImage
Ben Deen, Rebecca Saxe, Nancy Kanwisher
2020 J jnl
NeuroImage
Leyla Isik, Anna Mynick, Dimitrios Pantazis, Nancy Kanwisher
2019 J jnl
NeuroImage
Michael A. Cohen, Daniel D. Dilks, Kami Koldewyn, Sarah Weigelt, Jenelle Feather, Alexander J. E. Kell, Boris Keil, Bruce Fischl, Lilla Zöllei, Lawrence L. Wald, Rebecca Saxe, Nancy Kanwisher
2018 B conf
CogSci
Sarah Schwettmann, Jason Fischer, Josh Tenenbaum, Nancy Kanwisher
2018 J jnl
NeuroImage
Leyla Isik, Jedediah Singer, Joseph R. Madsen, Nancy Kanwisher, Gabriel Kreiman
2016 J jnl
CoRR
Daniel Harari, Tao Gao, Nancy Kanwisher, Joshua B. Tenenbaum, Shimon Ullman
2016 J jnl
NeuroImage
Frederik S. Kamps, Joshua B. Julian, Jonas Kubilius, Nancy Kanwisher, Daniel D. Dilks
2014 J jnl
NeuroImage
Anastasia Yendiki, Kami Koldewyn, Sita Kakunoori, Nancy Kanwisher, Bruce Fischl
2012 J jnl
NeuroImage
Joshua B. Julian, Evelina Fedorenko, Jason Webster, Nancy Kanwisher
2012 J jnl
NeuroImage
Danial Lashkari, Ramesh Sridharan, Edward Vul, Po-Jang Hsieh, Nancy Kanwisher, Polina Golland
2011 conf
MLINI
George H. Chen, Evelina Fedorenko, Nancy Kanwisher, Polina Golland
2011 J jnl
NeuroImage
David Pitcher, Daniel D. Dilks, Rebecca Saxe, Christina Triantafyllou, Nancy Kanwisher
2011 J jnl
Lang. Linguistics Compass
Evelina Fedorenko, Nancy Kanwisher
2011 J jnl
J. Cogn. Neurosci.
Evelina Fedorenko, Nancy Kanwisher
2010 J jnl
NeuroImage
Danial Lashkari, Ed Vul, Nancy Kanwisher, Polina Golland
2010 conf
CVPR Workshops
Danial Lashkari, Ramesh Sridharan, Ed Vul, Po-Jang Hsieh, Nancy Kanwisher, Polina Golland
2010 J jnl
J. Cogn. Neurosci.
Jia Liu, Alison Harris, Nancy Kanwisher
2009 J jnl
Lang. Linguistics Compass
Evelina Fedorenko, Nancy Kanwisher
2008 conf
MICCAI (1)
Danial Lashkari, Ed Vul, Nancy Kanwisher, Polina Golland
2006 J jnl
NeuroImage
Rebecca Saxe, Matthew Brett, Nancy Kanwisher
2003 J jnl
NeuroImage
Rebecca Saxe, Nancy Kanwisher
2002 conf
MICCAI (1)
Polina Golland, Bruce Fischl, Mona Spiridon, Nancy Kanwisher, Randy L. Buckner, Martha Elizabeth Shenton, Ron Kikinis, Anders M. Dale, W. Eric L. Grimson
CLAUDE.md
← Index CLAUDE.md markdown
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## Project Overview

REDB (RationalEdge Samples DB) is a malware analysis framework that extracts features from PE (Portable Executable) files and stores them in ClickHouse database for analysis. It provides a comprehensive set of extractors for analyzing binary samples including PE headers, imports, resources, signatures, and decompiled code.

## Common Commands

### Development Setup
```bash
source venv/bin/activate

# Install dependencies
pip install -r requirements.txt

# Run the main application
python start.py --path /path/to/samples --repo sample_repo --index_prefix redb
```

### Analysis Commands
```bash
# Process a single file
python start.py --path /path/to/binary --repo test --index_prefix redb

# Process from S3 storage
python start.py --s3 --repo malpedia --index_prefix redb

# Process from S3 storage but only a subset of a specific repository
python start.py --s3 --repo "vx-itw" --s3-notes "ITW.0138" --index_prefix redb

# Run only decompilation
python start.py --path /path/to/binary --repo test --index_prefix redb --decompile

# Run specific modules
python start.py --path /path/to/binary --repo test --index_prefix redb --modules "BasicPropertiesExtractor,PEFeaturesExtractor"

# Run as Nomad job (for containerized deployment)
python start.py --nomad-job
```

### Testing
There are no formal unit tests. Testing is done by running the extractors on sample files in the `test_files/` directory.

## Architecture Overview

### Core Components

1. **Ingestor (`redb/ingestor.py`)**: Main orchestrator that handles file processing, multiprocessing, and coordinates extractors
2. **Extractors (`redb/extractors/`)**: Modular analysis components that extract specific features
3. **Database Exporters (`redb/extractors/database_exporters.py`)**: Handle data export to ClickHouse
4. **Settings (`redb/settings/`)**: Configuration management for database connections

### Extractor Architecture

All extractors inherit from the base `Extractor` class and implement:
- `extract()`: Main analysis logic
- `prepare_export_data()`: Format data for database export
- `get_clickhouse_table()`: Return target table name

Available extractors:
- **General**: BasicPropertiesExtractor, HashExtractor, DIEExtractor, CAPAExtractor
- **PE-specific**: PEFeaturesExtractor, PEImportExtractor, PEResourceExtractor, PEOverlayExtractor, PESectionExtractor, PESignatureExtractor, PEDotNetExtractor, PEInconstistencyTestsExtractor, PEExtraFindings
- **ELF**: ELFFeaturesExtractor, ELFSegmentExtractor, ELFSectionExtractor, ELFDependencyExtractor, ELFSymbolExtractor, ELFImportExtractor, ELFExportExtractor, ELFRelocationExtractor, ELFNotesExtractor
- **Mach-O**: MachOFeaturesExtractor, MachOSegmentExtractor, MachOImportExtractor, MachOExportExtractor, MachODylibExtractor, MachOSignatureExtractor
- **APK**: APKFeaturesExtractor, APKManifestExtractor, APKPermissionsExtractor, APKSignatureExtractor, APKDexExtractor, APKResourceExtractor, APKNativeLibExtractor, APKInconsistencyTestsExtractor
- **Decompilation**: DecompileBinja, DecompileAPK

### Database Schema

The project uses a comprehensive ClickHouse schema defined in `redb/redb_schema.yml` with tables for:
- Basic properties (`redb_basic_properties`)
- PE features (`redb_pe_features`, `redb_pe_imports`, `redb_pe_sections`, etc.)
- Decompiled code (`code_binja_decompiled_functions_content`, `code_binja_decompiled_functions_references`)
- CAPA analysis (`redb_capa`, `redb_capa_capabilities`)

Full schema documentation is available in `docs/database_schema.md`.

### Processing Modes

1. **Analysis Mode**: Extracts features using selected modules
2. **Decompile Mode**: Uses Binary Ninja for code decompilation
3. **S3 Mode**: Fetches samples from S3 storage based on catalog queries
4. **Nomad Job Mode**: Processes single jobs using environment variables for containerized deployment

### Configuration

Environment variables are used for configuration:
- Database connection: `CLICKHOUSE_HOST`, `CLICKHOUSE_PORT`, `CLICKHOUSE_USER`, `CLICKHOUSE_PASSWORD`
- S3 storage: `S3_ENDPOINT`, `S3_ACCESS_KEY`, `S3_SECRET_KEY`
- Processing: `BATCH_SIZE`, `REDB_TIMEOUT`, `DECOMPILE_WORKER_TIMEOUT`
- Nomad jobs: `JOB_ID`, `S3_KEY`, `S3_BUCKET`, `WORKER_TYPE`, `CALLBACK_URL`, `ANALYSIS_MODULES`

## Important Implementation Details

### Multiprocessing
- Uses `spawn` method for multiprocessing to avoid memory issues
- Worker processes have timeout handlers to prevent hanging
- Supports both batch processing and streaming processing modes

### Memory Management
- Implements aggressive garbage collection between batches
- Monitors swap usage and restarts worker pools when needed
- Kills stuck processes automatically

### Error Handling
- Comprehensive logging with per-file context
- Graceful handling of corrupted or unsupported files
- Automatic retry logic for database operations

### Security Context
This is a defensive security tool for malware analysis. It processes potentially malicious files in a controlled environment to extract features for detection and analysis purposes.

## Development Notes

- The codebase is optimized for processing large batches of malware samples
- Extractors are designed to be modular and can be run individually or in combination
- Database schema supports both normalized and denormalized views for different query patterns
- S3 integration allows for scalable processing of large malware repositories