As part of the ENCODE 4 Project, we trained ChromBPNet models on 1,512 ENCODE DNAse-seq and ATAC-seq across 408 biosamples. Here, we provide all models for open-source use.
Modellquelle
Quellenbeschreibung
As part of the ENCODE 4 Project, we trained ChromBPNet models on 1,512 ENCODE DNAse-seq and ATAC-seq across 408 biosamples. Here, we provide all models for open-source use.
For more information about the models, see:
Quellen
1 QuelleVerifiziert 9. Aug.
Modellartefakte
15 Artefaktefold_0/model.bias_scaled.fold_0.ENCSR658UBE.h5
h5 · 2,56 MB · SHA-256 debe318f3d61…e57f · Hugging Face
Herunterladenfold_0/model.chrombpnet_nobias.fold_0.ENCSR658UBE.h5
h5 · 24,4 MB · SHA-256 557a3056848c…62fc · Hugging Face
HerunterladenQuellenauszüge
3 Auszügefold_0: Model of 5-fold cross-validation: Fold 0
model.chrombpnet.fold_0.encid.h5: full chrombpnet model that combines both bias and corrected model in .h5 formatmodel.chrombpnet_nobias.fold_0.encid.h5: bias-corrected accessibility model in .h5 format (Use for all biological discovery)model.bias_scaled.fold_0.encid.h5: bias model in .h5 formatmodel.chrombpnet.fold_0.encid.tar: full chrombpnet model that combines both bias and corrected model in SavedModel format. After being untarred, it results in a directory named "chrombpnet".model.chrombpnet_nobias.fold_0.encid.tar: bias-corrected accessibility model in SavedModel format (Use for all biological discovery). After being untarred, it results in a directory named "chrombpnet_wo_bias".model.bias_scaled.fold_0.encid.tar: bias model in SavedModel format. After being untarred, it results in a directory named "bias_model_scaled".logs.models.fold_0.encid: folder containing log files for training modelsfold_1: Model of 5-fold coss-validation: Fold 1fold_2: Model of 5-fold cross-validation: Fold 2fold_3: Model of 5-fold cross-validation: Fold 3fold_4: Model of 5-fold cross-validation: Fold 4(1) Use the code in python after appropriately defining model_in_h5_format and inputs.
(2) inputs is a one hot encoded sequence of shape (N,2114,4). Here N corresponds to the
number of tested sequences, 2114 is the input sequence length and 4 corresponds to [A,C,G,T].
import tensorflow as tf
from tensorflow.keras.utils import get_custom_objects
from tensorflow.keras.models import load_model
custom_objects={"tf": tf}
get_custom_objects().update(custom_objects)
model=load_model(model_in_h5_format,compile=False)
outputs = model(inputs)
The list outputs consists of two elements. The first element has a shape of (N, 1000) and
contains logit predictions for a 1000-base-pair output. The second element, with a shape of
(N, 1), contains logcount predictions. To transform these predictions into per-base signals,
follow the provided pseudo code lines below.
import numpy as np
def softmax(x, temp=1):
norm_x = x - np.mean(x,axis=1, keepdims=True)
return np.exp(temp*norm_x)/np.sum(np.exp(temp*norm_x), axis=1, keepdims=True)
predictions = softmax(outputs[0]) * (np.exp(outputs[1])-1)
(1) First untar the directory as follows tar -xvf model.tar.
(2) Use the code below in python after appropriately defining model_dir_untared and inputs.
(3) inputs is a one hot encoded sequence of shape (N,2114,4). Here N corresponds to the number
of tested sequences, 2114 is the input sequence length and 4 corresponds to ACGT.
import tensorflow as tf
model = tf.saved_model.load('model_dir_untared')
outputs = model.signatures['serving_default'](**{'sequence':inputs.astype('float32')})
The variable outputs represents a dictionary containing two key-value pairs. The first key
is logits_profile_predictions, holding a value with a shape of (N, 1000). This value corresponds
to logit predictions for a 1000-base-pair output. The second key, named `logcount_predictions``,
is associated with a value of shape (N, 1), representing logcount predictions. To transform these
predictions into per-base signals, utilize the provided pseudo code lines mentioned below.
import numpy as np
def softmax(x, temp=1):
norm_x = x - np.mean(x,axis=1, keepdims=True)
return np.exp(temp*norm_x)/np.sum(np.exp(temp*norm_x), axis=1, keepdims=True)
predictions = softmax(outputs["logits_profile_predictions"]) * (np.exp(outputs["logcount_predictions"])-1)
External data users may freely download, analyze and publish results based on any ENCODE data without restrictions.
Released under the ENCODE data-use policy. Please cite the ENCODE Project Consortium and the model software: ChromBPNet (Pampari et al., bioRxiv 2024).
fold_0/model.chrombpnet.fold_0.ENCSR658UBE.h5
h5 · 25,2 MB · SHA-256 55e1f7ad2012…5435 · Hugging Face
Herunterladenfold_1/model.bias_scaled.fold_1.ENCSR658UBE.h5
h5 · 2,56 MB · SHA-256 fe184e6cd27f…345f · Hugging Face
Herunterladenfold_1/model.chrombpnet_nobias.fold_1.ENCSR658UBE.h5
h5 · 24,4 MB · SHA-256 09ee1e299695…c6d4 · Hugging Face
Herunterladenfold_1/model.chrombpnet.fold_1.ENCSR658UBE.h5
h5 · 25,2 MB · SHA-256 7af51b856795…74f0 · Hugging Face
Herunterladenfold_2/model.bias_scaled.fold_2.ENCSR658UBE.h5
h5 · 2,56 MB · SHA-256 4a8d42b3e7fa…0301 · Hugging Face
Herunterladenfold_2/model.chrombpnet_nobias.fold_2.ENCSR658UBE.h5
h5 · 24,4 MB · SHA-256 625e1a8631c2…648f · Hugging Face
Herunterladenfold_2/model.chrombpnet.fold_2.ENCSR658UBE.h5
h5 · 25,2 MB · SHA-256 ecc41e4418b7…acaf · Hugging Face
Herunterladenfold_3/model.bias_scaled.fold_3.ENCSR658UBE.h5
h5 · 2,56 MB · SHA-256 5922998a2af2…e118 · Hugging Face
Herunterladenfold_3/model.chrombpnet_nobias.fold_3.ENCSR658UBE.h5
h5 · 24,4 MB · SHA-256 cbf6964d4f91…387c · Hugging Face
Herunterladenfold_3/model.chrombpnet.fold_3.ENCSR658UBE.h5
h5 · 25,2 MB · SHA-256 18294bc04ec1…a29f · Hugging Face
Herunterladenfold_4/model.bias_scaled.fold_4.ENCSR658UBE.h5
h5 · 2,56 MB · SHA-256 8a6f290d6934…cca7 · Hugging Face
Herunterladenfold_4/model.chrombpnet_nobias.fold_4.ENCSR658UBE.h5
h5 · 24,4 MB · SHA-256 a4a8e01934a0…bccb · Hugging Face
Herunterladenfold_4/model.chrombpnet.fold_4.ENCSR658UBE.h5
h5 · 25,2 MB · SHA-256 11092b4585fa…2d9d · Hugging Face
Herunterladen--- license: mit library_name: chrombpnet tags: - encode - chrombpnet - chromatin-accessibility - DNASE - lung - hg38 --- # ENCODE ChromBPNet Atlas As part of the ENCODE 4 Project, we trained ChromBPNet models on 1,512 ENCODE DNAse-seq and ATAC-seq across 408 biosamples. Here, we provide all models for open-source use. For more information about the models, see: - Main ENCODE 4 Paper - [A unified lexicon of predictive DNA sequence motifs from ENCODE transcription factor binding and chromatin accessibility assays](https://doi.org/10.5281/zenodo.17123347) (Yun, C. M. et al., Zenodo 2026) - [ChromBPNet: bias factorized, base-resolution deep learning models of chromatin accessibility reveal cis-regulatory sequence syntax, transcription factor footprints and regulatory variants](https://doi.org/10.1101/2024.12.25.630221) (Pampari, A. et al., bioRxiv 2024) ## ChromBPNet model: DNASE in left lung (ENCSR658UBE) - Model: ChromBPNet - Assay: DNASE-seq - Experiment: [ENCSR658UBE](https://www.encodeproject.org/experiments/ENCSR658UBE/) - Model annotation: [ENCSR773MZG](https://www.encodeproject.org/annotations/ENCSR773MZG/) - Biosample: left lung (Full name: Homo sapiens left lung tissue female child (16 years)) - Cell slim(s): None - Organ slim(s): lung - Developmental slim(s): endoderm - System slim(s): respiratory-system - Assembly: hg38 ## Directory structure - `fold_0`: Model of 5-fold cross-validation: Fold 0 - `model.chrombpnet.fold_0.encid.h5`: full chrombpnet model that combines both bias and corrected model in .h5 format - `model.chrombpnet_nobias.fold_0.encid.h5`: bias-corrected accessibility model in .h5 format (Use for all biological discovery) - `model.bias_scaled.fold_0.encid.h5`: bias model in .h5 format - `model.chrombpnet.fold_0.encid.tar`: full chrombpnet model that combines both bias and corrected model in SavedModel format. After being untarred, it results in a directory named "chrombpnet". - `model.chrombpnet_nobias.fold_0.encid.tar`: bias-corrected accessibility model in SavedModel format (Use for all biological discovery). After being untarred, it results in a directory named "chrombpnet_wo_bias". - `model.bias_scaled.fold_0.encid.tar`: bias model in SavedModel format. After being untarred, it results in a directory named "bias_model_scaled". - `logs.models.fold_0.encid`: folder containing log files for training models - `fold_...
--- license: mit library_name: chrombpnet tags: - encode - chrombpnet - chromatin-accessibility - DNASE - lung - hg38 --- # ENCODE ChromBPNet Atlas As part of the ENCODE 4 Project, we trained ChromBPNet models on 1,512 ENCODE DNAse-seq and ATAC-seq across 408 biosamples. Here, we provide all models for open-source use. For more information about the models, see: - Main ENCODE 4 Paper - [A unified lexicon of predictive DNA sequence motifs from ENCODE transcription factor binding and chromatin accessibility assays](https://doi.org/10.5281/zenodo.17123347) (Deshpande et al., Zenodo 2025) - [ChromBPNet: bias factorized, base-resolution deep learning models of chromatin accessibility reveal cis-regulatory sequence syntax, transcription factor footprints and regulatory variants](https://doi.org/10.1101/2024.12.25.630221) (Pampari et al., bioRxiv 2024) ## ChromBPNet model: DNASE in left lung (ENCSR658UBE) - Model: ChromBPNet - Assay: DNASE-seq - Experiment: [ENCSR658UBE](https://www.encodeproject.org/experiments/ENCSR658UBE/) - Model annotation: [ENCSR773MZG](https://www.encodeproject.org/annotations/ENCSR773MZG/) - Biosample: left lung (Full name: Homo sapiens left lung tissue female child (16 years)) - Cell slim(s): None - Organ slim(s): lung - Developmental slim(s): endoderm - System slim(s): respiratory-system - Assembly: hg38 ## Directory structure - `fold_0`: Model of 5-fold cross-validation: Fold 0 - `model.chrombpnet.fold_0.encid.h5`: full chrombpnet model that combines both bias and corrected model in .h5 format - `model.chrombpnet_nobias.fold_0.encid.h5`: bias-corrected accessibility model in .h5 format (Use for all biological discovery) - `model.bias_scaled.fold_0.encid.h5`: bias model in .h5 format - `model.chrombpnet.fold_0.encid.tar`: full chrombpnet model that combines both bias and corrected model in SavedModel format. After being untarred, it results in a directory named "chrombpnet". - `model.chrombpnet_nobias.fold_0.encid.tar`: bias-corrected accessibility model in SavedModel format (Use for all biological discovery). After being untarred, it results in a directory named "chrombpnet_wo_bias". - `model.bias_scaled.fold_0.encid.tar`: bias model in SavedModel format. After being untarred, it results in a directory named "bias_model_scaled". - `logs.models.fold_0.encid`: folder containing log files for training models - `fold_1`: M...
Source context: 0 downloads · 0 likes · Library chrombpnet · Repo kundajelab/encode-chrombpnet-DNASE-ENCSR658UBE-ENCSR773MZG