justIN           Dashboard       Workflows       Jobs       AWT       Sites       Storages       Docs       Login

Workflow 7081, Stage 1

Priority50
Processors1
Wall seconds80000
Image/cvmfs/singularity.opensciencegrid.org/fermilab/fnal-wn-sl7:latest
RSS bytes3144679424 (2999 MiB)
Max distance for inputs30.0
Enabled input RSEs CERN_PDUNE_EOS, DUNE_CA_SFU, DUNE_CERN_EOS, DUNE_ES_PIC, DUNE_FR_CCIN2P3_DISK, DUNE_IN_TIFR, DUNE_IT_INFN_CNAF, DUNE_UK_GLASGOW, DUNE_UK_LANCASTER_CEPH, DUNE_UK_MANCHESTER_CEPH, DUNE_US_BNL_SDCC, DUNE_US_FNAL_DISK_STAGE, FNAL_DCACHE, FNAL_DCACHE_STAGING, FNAL_DCACHE_TEST, MONTECARLO, NIKHEF, PRAGUE, QMUL, RAL-PP, RAL_ECHO, SURFSARA, T3_US_NERSC
Enabled output RSEs CERN_PDUNE_EOS, DUNE_CA_SFU, DUNE_CERN_EOS, DUNE_ES_PIC, DUNE_FR_CCIN2P3_DISK, DUNE_IN_TIFR, DUNE_IT_INFN_CNAF, DUNE_UK_GLASGOW, DUNE_UK_LANCASTER_CEPH, DUNE_UK_MANCHESTER_CEPH, DUNE_US_BNL_SDCC, DUNE_US_FNAL_DISK_STAGE, FNAL_DCACHE, FNAL_DCACHE_STAGING, FNAL_DCACHE_TEST, NIKHEF, PRAGUE, QMUL, RAL-PP, RAL_ECHO, SURFSARA, T3_US_NERSC
Enabled sites BR_CBPF, CA_SFU, CA_Victoria, CERN, CH_UNIBE-LHEP, ES_CIEMAT, ES_PIC, FR_CCIN2P3, IN_TIFR, IT_CNAF, NL_SURFsara, UK_Bristol, UK_Brunel, UK_Durham, UK_Edinburgh, UK_Lancaster, UK_Manchester, UK_Oxford, UK_QMUL, UK_RAL-PPD, UK_RAL-Tier1, UK_Sheffield, US_Caltech, US_Colorado, US_FNAL-FermiGrid, US_FNAL-T1, US_Michigan, US_MIT, US_Nebraska, US_NotreDame, US_PuertoRico, US_SU-ITS, US_Swan, US_UChicago, US_UConn-HPC, US_UCSD, US_Wisconsin
Scopeusertests
Events for this stage

Output patterns

 DestinationPatternLifetimeFor next stageRSE expression
1Rucio usertests:calcuttj_test_convert_unet-w7081s1p1*.pt604800False

Environment variables

NameValue
DOCOPY1
NFILES5
UNET_DIR/cvmfs/fifeuser3.opensciencegrid.org/sw/dune/b5765e0d70ac8df928f55219574a442063a5eb63

File states

Total filesFindingUnallocatedAllocatedOutputtingProcessedNot foundFailed
100900100

Job states

TotalSubmittedStartedProcessingOutputtingFinishedNotusedAbortedStalledJobscript errorOutputting failedNone processed
44000034082000
Files processed000.10.10.20.20.30.30.40.40.50.50.60.60.70.70.80.80.90.911May-21 14:00May-21 15:00May-21 16:00Files processedBin start timesNumber per binUK_RAL-PPD
Replicas per RSE4480.1462511658797201.460510492318042380.0339.32279.8537488341203266.5394895076821294.81051049231803172.10621293360261347.460510492318133.85374883412032Replicas per RSEDUNE_UK_MANCHESTER_CEPH (40…DUNE_UK_MANCHESTER_CEPH (40%)RAL_ECHO (20%)RAL-PP (20%)DUNE_FR_CCIN2P3_DISK (10%)DUNE_US_FNAL_DISK_STAGE (10…DUNE_US_FNAL_DISK_STAGE (10%)

RSEs used

NameInputsOutputs
DUNE_UK_MANCHESTER_CEPH120
RAL_ECHO60
RAL-PP65
DUNE_US_FNAL_DISK_STAGE40
DUNE_FR_CCIN2P3_DISK10

Stats of processed input files as CSV or JSON, and of uploaded output files as CSV or JSON (up to 10000 files included)

File reset events, by site

SiteAllocatedOutputting
UK_RAL-PPD110
UK_Edinburgh80
FR_CCIN2P350
US_Colorado40

Jobscript

#!/bin/bash
#
if [ -z $UNET_DIR ]; then
        echo "fatal must provide UNET_DIR env var"
        exit 1
fi

source /cvmfs/dune.opensciencegrid.org/products/dune/setup_dune.sh
setup python v3_9_15
setup xrootd v5_5_5a -q e26:p3915:prof

if [ -z ${JUSTIN_PROCESSORS} ]; then
  JUSTIN_PROCESSORS=1
fi

echo "Justin processors: ${JUSTIN_PROCESSORS}"

export TF_NUM_THREADS=${JUSTIN_PROCESSORS}   
export OPENBLAS_NUM_THREADS=${JUSTIN_PROCESSORS} 
export JULIA_NUM_THREADS=${JUSTIN_PROCESSORS} 
export MKL_NUM_THREADS=${JUSTIN_PROCESSORS} 
export NUMEXPR_NUM_THREADS=${JUSTIN_PROCESSORS} 
export OMP_NUM_THREADS=${JUSTIN_PROCESSORS}  

NFILES=${NFILES:-1}
files=()
nfiles=0
#if [ $NFILES -eq 1 ]; then
#  echo "Will use justin-get-file"
#  DID_PFN_RSE=`$JUSTIN_PATH/justin-get-file`
#  if [ "${DID_PFN_RSE}" == "" ] ; then
#    echo "Could not get file"
#    exit 0
#  fi
#  export pfn=`echo ${DID_PFN_RSE} | cut -f2 -d' '`
#  export did=`echo ${DID_PFN_RSE} | cut -f1 -d' '`
#
#  if [ -z ${TESTFILE} ]; then
#    files+=($pfn)
#  else
#    files+=(${TESTFILE})
#  fi
#
if [ -n "${TESTFILE}" ]; then
  echo "Using testfile"
  files+=(${TESTFILE})
  nfiles=1
else
  nfiles=0
  for i in `seq 1 ${NFILES:-1}`; do
    echo $i
    DID_PFN_RSE=`$JUSTIN_PATH/justin-get-file`

    if [ "${DID_PFN_RSE}" == "" ] ; then
      echo "Could not get file -- exiting loop"
      break
    fi

    echo "did_pfn_rse: ${DID_PFN_RSE}"

    THISFILE=`echo ${DID_PFN_RSE} | cut -f2 -d' '`

    echo $THISFILE
    files+=(${THISFILE})
    nfiles=$(( nfiles + 1 ))
  done
fi

echo "Got $nfiles Files"
echo ${files[@]}

if [ $nfiles -eq 0 ]; then
  echo "Got no files. Exiting safely"
fi


echo "Justin specific env vars"
env | grep JUSTIN
now=$(date -u +"%Y%m%dT%H%M%SZ")
jobid=`echo "${JUSTIN_JOBSUB_ID:-1}" | cut -f1 -d'@' | sed -e "s/\./_/"`
stageid=${JUSTIN_STAGE_ID:-1}


if [ -z ${TORCHDIR} ]; then
  echo "installing torch"
  pip install --user torch h5py
  exitcode=$?
  if [ $exitcode -ne 0 ]; then
    echo "Error installing torch. Exiting with ${exitcode}"
    exit $exitcode
  fi
else
  echo "using venv"
  source ${TORCHDIR}/bin/activate
fi

if [ -z ${DOCOPY} ]; then 
  echo "Not copyign"
  input=${files[@]}
else
  for f in ${files[@]}; do
    echo "Copying $f"
    xrdcp $f .
  done
  input=`ls *h5`
fi 

echo "input ${input[@]}"

## Run convert on the output
for f in $input; do
  echo $f
  LD_PRELOAD=$XROOTD_LIB/libXrdPosixPreload.so python $UNET_DIR/convert_to_pt.py $f #$files #$pfn #extracted*h5
  exitcode=$?
  if [ $exitcode -ne 0 ]; then
    echo "Error converting to torch. Exiting with ${exitcode}"
    exit $exitcode
  fi
done

#echo "$pfn" > justin-processed-pfns.txt
for i in ${files[@]}; do
  echo "${i}" >> justin-processed-pfns.txt
done
justIN time: 2025-05-22 23:07:28 UTC       justIN version: 01.03.01
<