***Note:***

Publications making use of the DREAMS DR1 data should cite: 
    Yang, Zang, Valdes et al. (2026) [DOI:10.3847/1538-4365/ae8fd4].
See the paper for detailed descriptions of the processing and derived parameters.

DREAMS is based on observations made at NSF Cerro Tololo Inter-American Observatory, NSF NOIRLab (NOIRLab Prop. ID 2025A-806294, PI: W.Zang; NOIRLab Prop. ID: 2025B-560332 PIs: H. Yang and W. Zang), which is managed by the Association of Universities for Research in Astronomy (AURA) under a cooperative agreement with the U.S. National Science Foundation. This project used data obtained with the Dark Energy Camera (DECam), which was constructed by the Dark Energy Survey (DES) collaboration. For the full acknowledgement, please see acknowledgement.tex.

# DREAMS DR1 data

The file structure is as follows:
``` text
DREAMS_DR1/
├── src/                # bulid-in scripts
├── data/               
│   ├── catalog/        # catalog data
│   │   └── group/      # support star group catalogs
│   └── lightcurve/     # all light curves
│       ├── lc_xxxx.h5  # individual HDF5 light curve files
│       └── ...         # individual HDF5 light curve files
├── output/             # output of read_lc.py will be here
├── search.py           # script to search stars in the catalog
├── read_lc.py          # script to read light curves for individual stars
├── eventfinder.py      # simplified example of candidate search
├── environment.yml     # for dependency environment set up
├── README.md
├── acknowledgement.tex # full acknowledgement.
└── files.md5           # MD5 values for complete verification
```

The full light-curve archive is large, so the files are stored in a customized HDF5 layout for efficient access:

- Entire light-curve database: about 2.2 TB
- Stored in 7808 HDF5 (`.h5`) files
- z band: about 3904 files, each about 500 MB
- r band: about 3904 files, each about 50 MB

Loading the full archive into memory at once is unrealistic. This repository provides lightweight tools for working with the DREAMS DR1 light-curve archive:

- search.py for catalog queries by star ID, cone search, or box search
- read_lc.py for loading one star's light curves and related metadata on demand
- eventfinder.py as a short example of how to scan a support catalog for candidate events


## Environment Setup

The examples below assume that conda is already installed.

Create the recommended environment from environment.yml:

```bash
conda env create -f environment.yml
conda activate dreams
```

Main runtime dependencies are:

- numpy
- astropy
- h5py
- healpy
- matplotlib

The code expects the data directories to remain in the repository layout used here, especially:

- data/catalog/DREAMS_star_catalog_v1.h5
- data/catalog/group/
- data/lightcurve/


## Catalog Search with search.py

search.py exposes both a command-line interface and a Python class wrapper.

### Command Line Usage

Search by one or more global IDs:

```bash
python search.py id 46176216 47807210 49665382
```

Cone search around a sky position in decimal degrees: (radius in arcseconds)

```bash
python search.py circle 267.63463 -30.00255 2.0
```

Cone search also accepts sexagesimal coordinates:

```bash
python search.py circle 17:50:32.311 -30:00:09.18 2.0
```

Box search in a rectangular region:

```bash
python search.py box 267.634 267.635 -30.003 -30.002
```

Optional sorting is available:

```bash
python search.py circle 267.63463 -30.00255 2.0 --sort global_id
python search.py box 267.634 267.635 -30.003 -30.002 --sort ra
python search.py id 46176216 47807210 --sort none
```

The script prints matching catalog rows, including:

- global_id
- ra, dec
- catalog magnitude and magnitude error
- stamp names and internal IDs
- separation for cone searches

### Python Usage

If you want a small wrapper around the catalog database, import CatalogSearch from search.py:

```python
from search import CatalogSearch

catalog = CatalogSearch()

results_id = catalog.search_id([46176216, 47807210])
results_circle = catalog.search_circle(267.63463, -30.00255, 2.0)
results_box = catalog.search_box(267.634, 267.635, -30.003, -30.002)
```

You can also call the lower-level catalog database directly.

This is possible, but not recommended for workflows because the low-level API may change in future versions.

```python
from pathlib import Path
from src.CatalogDatabase import CatalogDatabase

cat_dir = Path('data/catalog')
catalogdb = CatalogDatabase(cat_dir / 'DREAMS_star_catalog_v1.h5')

results = catalogdb.search_id([46176216, 47807210])
results = catalogdb.search_circle(267.63463, -30.00255, 2.0)
results = catalogdb.search_box(267.634, 267.635, -30.003, -30.002)
```

## Light-Curve Reading with read_lc.py

read_lc.py is the main access point for retrieving per-star light curves from the sharded HDF5 archive.

### Command Line Usage

The built-in command-line entry point is best treated as a quick read test for a single star ID:

```bash
python read_lc.py 19178854
python read_lc.py 19178854 --version DR1.5
python read_lc.py 19178854 --version DR1
```

This interface loads the requested star through the internal reader and prints verbose progress. For most analysis workflows, importing read_lc in Python is more practical.

### Python Usage

Minimal usage:

```python
from read_lc import read_lc

light_curves = read_lc(19178854)
```

Common optional arguments:

```python
light_curves = read_lc(
    19178854,
    bands=['r', 'z'],
    read_cmd=True,
    read_image=False,
    version='DR1',
    verbose=True,
)
```

The return value is a list of dictionaries. Each dictionary corresponds to one available stamp and band combination for the requested star.

### Returned Fields

Core metadata fields:

- id (global)
- band
- stamp
- has_data
- mag_zero
- ra, dec
- cat_mag, cat_merr
- flux_base, ferr_base
- mag_base, merr_base
- ndata
- x, y

Image-related fields, only when read_image=True:

- ref_image: reference image as a 2D array
- ref_mask: reference mask as a 2D array

CMD-related fields, which contains all field stars on the reference image, only when read_cmd=True:

- base_id_all (internal)
- base_mag_all
- base_merr_all
- base_x_all
- base_y_all
- base_ra_all
- base_dec_all

Light-curve time-series fields:

- image_id
- mjd
- flux, ferr
- mag, merr
- fwhm
- sky
- qirr
- airmass
- stddev
- sigma_subt
- sigma_phot
- flag

Quality flag meaning:

- 0: good
- 1: bad DIA
- 2: bad PSF photometry
- 3: both DIA and PSF photometry are bad
- 4: out of or near the CCD boundary

### Example: Inspect One Star

```python
from read_lc import read_lc

light_curves = read_lc(19178854, read_cmd=True)

for lc in light_curves:
    if not lc['has_data']:
        continue
    print(lc['stamp'], lc['band'], lc['ndata'])
    print(lc['mjd'][:5], lc['mag'][:5])
```


## Support for Batch Access

The catalog database is only used to find stars efficiently, while the light curves stay sharded across many HDF5 files. For large-scale searches, it is often better to begin with a support group catalog such as:

```text
data/catalog/group/stargroup_D01.N8.0703_D02.S25.0303.cat
```

These files contain compact per-group star lists. A typical file starts with columns like:

```text
# global_id        ra          dec    mag  merr       stamp1 internal_id1 x1      y1   mag1 merr1       stamp2 internal_id2 x2      y2   mag2 merr2
01563413 267.60149403 -29.14035967 15.187 0.003  D01.N8.0703      31 377.240 373.830 15.187 0.003 D02.S25.0303       8 351.890  53.570 15.240 0.005
01563414 267.60176112 -29.13840584 15.278 0.004  D01.N8.0703      32 350.450 376.700 15.278 0.004 D02.S25.0303       9 324.980  56.530 15.347 0.005
01563415 267.60446190 -29.12274680 15.559 0.003  D01.N8.0703      34 135.660 406.420 15.559 0.003 D02.S25.0303      12 109.980  86.410 15.630 0.004
```

This makes it practical to loop over a few thousand stars in one group, which are typically stored in the same small set of HDF5 files (2~4), instead of attempting a blind scan over the full archive and opening unrelated files across the entire database.


## Event Search Example with eventfinder.py

eventfinder.py is a short example script that demonstrates a scalable search pattern. It is not a complete event-detection pipeline.

The recommended starting point is a support group catalog under data/catalog/group, for example:

```text
data/catalog/group/stargroup_D01.N8.0304_D02.S26.0704.cat
```

The basic idea is:

1. Load one support catalog.
2. Reuse the corresponding HDF5 datasets for that stamp combination.
3. Loop over stars in the group.
4. Call your own signal-search function.
5. Save or inspect diagnostic plots for candidates.

Minimal pattern:

```python
import numpy as np
from read_lc import read_lc

catalog = np.loadtxt(
    'data/catalog/group/stargroup_D01.N8.0304_D02.S26.0704.cat',
    dtype=str,
)

for star_id in catalog[:, 0]:
    light_curves = read_lc(int(star_id))

    has_signal, signal = signal_search(light_curves)

    if has_signal:
        plot_signal()
```

The provided eventfinder.py goes one step further by:

- loading reusable datasets for a chosen stamp combination
- parsing one support catalog into per-star metadata
- applying a toy event_search function
- calling read_lc again with read_cmd=True for candidate visualization

This pattern is much more efficient than repeatedly opening unrelated HDF5 files across the full archive.


## Practical Notes

- Prefer starting from search.py or a support group catalog before reading large numbers of light curves.
- A single star may return multiple light-curve dictionaries if it is covered by more than one stamp and band.
- read_image=True and read_cmd=True increase I/O and memory use and can be significantly slower, so they are best used after candidate selection rather than for all stars.


## Known Issues
- In DR1, light curves for stars near the CCD edge, as well as saturated stars, often show strong flux-seeing or flux-sky correlations. Users may need additional checks to filter these cases. This is expected to improve in future data releases.
