bone_2026/QWEN.md

194 lines
5.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Bone Quality Assessment Project
## Project Overview
Medical AI service for automated assessment of DXA (bone densitometry) study quality. The system analyzes DICOM files and evaluates quality based on standard criteria.
### Core Purpose (Hackathon)
- Analyze DICOM densitometry studies
- Determine anatomical region (spine/hip)
- Binary classification: quality (OK/violation)
- Output results in XLSX/CSV format per requirements
### Tech Stack
| Component | Technology |
|-----------|------------|
| Backend | Python 3.10, FastAPI, Uvicorn |
| ML/Deep Learning | PyTorch, torchvision (ResNet18) |
| Image Processing | PIL, OpenCV, pydicom |
| Data Handling | pandas, openpyxl |
| Containerization | Docker |
---
## Project Structure
```
bone_2026/
├── src/
│ ├── main.py # FastAPI app (DXA mode)
│ ├── run.py # Server runner
│ ├── dxa/ # DXA Quality module
│ │ ├── dataset.py # DXADataset class
│ │ ├── model.py # ResNet18 classifier
│ │ ├── train.py # Training script
│ │ ├── inference.py # Batch inference
│ │ └── __init__.py
│ ├── api/ # REST endpoints
│ ├── core/ # Orchestrator
│ ├── quality/ # Quality scoring
│ ├── segmentators/ # Segmentation models
│ └── classifiers/ # Classification models
├── models/
│ └── dxa_model.pth # Trained DXA classifier
├── dataset_hack/ # DICOM datasets
│ ├── Для теста/ # Test data (3 files)
│ └── НД_для_обучения/ # Training data (100 studies, 499 DICOMs)
│ └── разметка.xlsx # Annotation file
├── requirements.txt
├── Dockerfile
├── run.sh # Main entry script
└── README.md
```
---
## DXA Module (`src/dxa/`)
### Dataset (`dataset.py`)
- Loads DICOM files from studies
- Parses annotation Excel file
- Maps anatomical regions: spine, hip_right, hip_left
- Quality labels: 0 (OK), 1 (violation)
- Total: ~1433 samples (train: 1146, val: 287)
### Model (`model.py`)
- Architecture: ResNet18 (pretrained on ImageNet)
- Task: Binary classification (quality OK vs violation)
- Input: 224x224 RGB images
- Output: class probabilities
### Training (`train.py`)
```bash
python src/dxa/train.py --epochs 10 --batch-size 16
```
### Inference (`inference.py`)
```bash
python src/dxa/inference.py \
--input-path dataset_hack/Для\ теста \
--output-path results.xlsx \
--model-path models/dxa_model.pth
```
---
## Running the Project
### Training
```bash
# Option 1: Direct Python
python src/dxa/train.py --epochs 10
# Option 2: Via run.sh
bash run.sh train
```
### Inference
```bash
# Single file
python src/dxa/inference.py --input-path file.dcm --output-path result.xlsx
# Directory (batch)
python src/dxa/inference.py --input-path dataset_hack/Для\ теста --output-path results.xlsx
```
### API Server
```bash
python -m uvicorn src.main:app --host 0.0.0.0 --port 8000
```
---
## Output Format (per Hackathon Requirements)
| Column | Description |
|--------|-------------|
| path_to_study | Path to study directory |
| study_uid | StudyInstanceUID from DICOM |
| image_uid | SOPInstanceUID from DICOM |
| anatomical_region | spine / hip |
| quality_class | 0 (OK), 1 (violation) |
| violation_type | Type of violation (if any) |
| processing_status | Success / Failure |
| time_of_processing | Processing time (seconds) |
---
## Model Performance
```
Training data: 1146 samples
Validation data: 287 samples
Training (10 epochs):
- Best F1: ~0.27 (imbalanced classes: ~65% OK, ~35% violation)
- Accuracy: ~84%
Note: Need more epochs (50+) and class balancing for production
```
---
## Annotation Format
The annotation Excel (`разметка.xlsx`) contains:
- Study UID
- Spine columns: укладка, ось, артефакты
- Hip columns: позиция, ROI (left/right)
- Total columns: итого
---
## Development Conventions
### Code Style
- Follow existing patterns in src/
- Type hints where appropriate
- Minimal comments (only for context)
### Key Components
- **DXADataset**: Handles DICOM loading + annotation parsing
- **DXAQualityClassifier**: ResNet18-based classifier
- **process_dicom_files**: Batch inference with XLSX output
### Dependencies
All in `requirements.txt`:
- `torch`, `torchvision` - Deep learning
- `pydicom` - DICOM handling
- `pandas`, `openpyxl` - Data/Excel
- `fastapi`, `uvicorn` - Web framework
- `Pillow`, `opencv-python-headless` - Image processing
---
## Docker
```bash
# Build
docker build -t dxa-quality .
# Run
docker run -v /data:/data -p 8000:8000 dxa-quality
```
---
## Notes
- This is a **hackathon project** for DXA quality assessment
- Model trained on limited data (100 studies)
- Binary classification (quality OK / violation)
- Anatomical region detection via image size heuristic
- Output format matches hackathon requirements (XLSX/CSV)