bone_2026/QWEN.md

247 lines
7.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Bone Quality Assessment Project
## Project Overview
Medical AI service for automated assessment of DXA (bone densitometry) study quality. The system analyzes DICOM files and evaluates quality based on standard criteria.
### Core Purpose (Hackathon)
- Analyze DICOM densitometry studies
- Determine anatomical region (spine/hip)
- Binary classification: quality (OK/violation)
- Detailed violation type detection (motion, artifacts, position, ROI)
- Output results in XLSX/CSV format per requirements
### Tech Stack
| Component | Technology |
|-----------|------------|
| Backend | Python 3.10, FastAPI, Uvicorn |
| ML/Deep Learning | PyTorch, torchvision (ResNet18) |
| Image Processing | PIL, OpenCV, pydicom, scipy |
| Data Handling | pandas, openpyxl |
| Containerization | Docker |
---
## Project Structure
```
bone_2026/
├── src/
│ ├── main.py # FastAPI app (DXA mode)
│ ├── run.py # Server runner
│ ├── dxa/ # DXA Quality module
│ │ ├── dataset.py # DXADataset class
│ │ ├── model.py # ResNet18 classifier
│ │ ├── train.py # Training script
│ │ ├── inference.py # Batch inference
│ │ └── __init__.py
│ ├── api/ # REST endpoints & schemas
│ │ ├── endpoints.py # Original endpoints
│ │ └── static/ # Web UI
│ │ ├── index.html
│ │ └── js/dxa-app.js
│ ├── core/ # Orchestrator
│ ├── quality/ # Quality scoring
│ │ ├── quality_scorer.py # Base scorer
│ │ ├── detailed_assessment.py # NEW: Detailed assessment
│ │ ├── position_validator.py
│ │ └── artifact_detector.py
│ └── segmentators/ # Segmentation models
├── models/
│ └── dxa_model.pth # Trained DXA classifier
├── dataset_hack/ # DICOM datasets
│ ├── Для теста/ # Test data
│ └── НД_для_обучения/ # Training data
├── requirements.txt
├── Dockerfile
├── run.sh
└── README.md
```
---
## Implemented Features (Current)
### ✅ API Endpoints
| Method | Endpoint | Description |
|--------|----------|-------------|
| GET | `/` | Web interface |
| GET | `/api/v1/health` | Health check |
| POST | `/api/v1/analyze` | Basic analysis |
| POST | `/api/v1/analyze/detailed` | **NEW: Detailed analysis with metrics** |
| POST | `/api/v1/analyze/sr` | **NEW: DICOM SR report** |
| POST | `/api/v1/batch` | Batch processing |
| POST | `/api/v1/export` | Export to XLSX |
### ✅ Detailed Assessment (`src/quality/detailed_assessment.py`)
- **Motion detection**: Laplacian variance, FFT blur analysis
- **Artifact detection**: Metal, implants, cement, calcifications
- **Spine completeness**: Vertebrae count, alignment, spacing
- **Hip completeness**: Full visibility, aspect ratio
- **Hip rotation**: Major axis angle calculation
- **ROI validation**: Boundary check, margin, size
- **Violation types**: correct, position_error, artifact_motion, artifact_other, labeling_error, incomplete_view, roi_error, rotation
### ✅ Web Interface
- Drag-and-drop DICOM upload
- Table with results (filter, sort, search)
- **NEW: Detail panel** - click on row to see:
- Violation type
- Reason (human-readable)
- Motion metrics
- Artifact detection
- ROI validation
---
## DXA Module (`src/dxa/`)
### Dataset (`dataset.py`)
- Loads DICOM files from studies
- Parses annotation Excel file
- **Automatically detects anatomical region from image content**
- Maps regions: spine, hip_right, hip_left
- Quality labels: 0 (OK), 1 (violation)
### Model (`model.py`)
- Architecture: ResNet18 (pretrained on ImageNet)
- Task: Binary classification (quality OK vs violation)
- Input: 224x224 RGB images
- Output: class probabilities
### Training (`train.py`)
```bash
python src/dxa/train.py --epochs 10 --batch-size 16
```
### Inference (`inference.py`)
```bash
python src/dxa/inference.py \
--input-path dataset_hack/Для\ теста \
--output-path results.xlsx \
--model-path models/dxa_model.pth
```
---
## Running the Project
### Training
```bash
python src/dxa/train.py --epochs 10
```
### Inference
```bash
python src/dxa/inference.py --input-path file.dcm --output-path result.xlsx
```
### API Server
```bash
python -m uvicorn src.main:app --host 0.0.0.0 --port 8000
```
Web UI: http://localhost:8000
---
## Output Format (per Hackathon Requirements)
### Basic Output
| Column | Description |
|--------|-------------|
| path_to_study | Path to study directory |
| study_uid | StudyInstanceUID from DICOM |
| image_uid | SOPInstanceUID from DICOM |
| anatomical_region | spine / hip_left / hip_right / hip |
| quality_class | 0 (OK), 1 (violation) |
| violation_type | Type of violation (if any) |
| processing_status | Success / Failure |
| time_of_processing | Processing time (seconds) |
### Detailed Output (/api/v1/analyze/detailed)
```json
{
"anatomical_region": "spine",
"quality_class": 1,
"quality_label": "Violation detected",
"violation_type": "artifact_motion",
"reason": "Обнаружен артефакт движения (размытие)",
"confidence": 0.85,
"confidence_per_class": {
"correct": 0.15,
"violation": 0.85
},
"view_quality": "full",
"metrics": {
"motion": { "motion_detected": true, "severity": "HIGH" },
"artifacts": { "any_detected": true, "metal_detected": false },
"roi_check": { "valid": true }
},
"overall_quality": "POOR",
"severity": "HIGH"
}
```
---
## Anatomical Region Detection
The system automatically determines the anatomical region from the DICOM image content:
### Algorithm (`src/dxa/inference.py`)
1. **Spine vs Hip** - by bright region shape:
- Extract 95th percentile threshold
- Calculate bounding box aspect ratio
- Spine: bbox_aspect < 1.5 (more square)
- Hip: bbox_aspect > 1.5 (vertically elongated)
2. **Hip Left vs Right** - by brightness asymmetry:
- Calculate left/right bright pixel ratio
- hip_left: L/R ratio < 0.7 (left side brighter)
- hip_right: L/R ratio > 1.3 (right side brighter)
- hip: unclear (fallback)
---
## Known Issues & Limitations
1. **Model training** - Needs retraining with new violation types
2. **Dataset size** - Currently ~100 studies, needs 500+
3. **Segmentation** - Uses simple threshold, needs proper model
4. **F1 score** - Currently ~0.27, needs improvement with weighted loss
5. **Heatmap visualization** - Not implemented (requires model retraining)
---
## Docker
```bash
docker build -t dxa-quality .
docker run -v /data:/data -p 8000:8000 dxa-quality
```
---
## Development Notes
### Code Style
- Follow existing patterns in src/
- Type hints where appropriate
- Minimal comments (only for context)
### Key Components
- **DXADataset**: Handles DICOM loading + annotation parsing
- **DXAQualityClassifier**: ResNet18-based classifier
- **process_dicom_files**: Batch inference with XLSX output
- **generate_quality_report**: Detailed assessment with metrics
### Dependencies
All in `requirements.txt`:
- `torch`, `torchvision` - Deep learning
- `pydicom` - DICOM handling
- `pandas`, `openpyxl` - Data/Excel
- `fastapi`, `uvicorn` - Web framework
- `Pillow`, `opencv-python-headless` - Image processing
- `scipy` - Image analysis (blur, artifacts)