justIN           Dashboard       Workflows       Jobs       AWT       Sites       Storages       Docs       Login

Jobsub ID 262644.0@justin-prod-sched01.dune.hep.ac.uk

Jobsub ID262644.0@justin-prod-sched01.dune.hep.ac.uk
Workflow TestingYes
Workflow ID1
Stage ID1
User nameamcnab@fnal.gov
HTCondor Groupgroup_dune.prod_mcsim
RequestedProcessors1
RSS bytes1073741824 (1024 MiB)
Wall seconds limit3600 (1 hours)
Submitted time2024-09-26 01:18:00
SiteUS_UCSD
EntryCMSHTPC_T2_US_UCSD_gw7
Last heartbeat2024-09-26 02:25:38
From worker nodeHostnamesdsc-66.t2.ucsd.edu
cpuinfoIntel(R) Xeon(R) CPU E5-2680 v4 @ 2.40GHz
OS releaseScientific Linux release 7.9 (Nitrogen)
Processors1
RSS bytes1073741824 (1024 MiB)
Wall seconds limit171000 (47 hours)
Inner Apptainer?True
Job statefinished
Allocator namejustin-allocator-pro.dune.hep.ac.uk
Started2024-09-26 01:23:59
Input files
JobscriptExit code0
Real time1h (3687s)
CPU time0m (17s = 0%)
Outputting started2024-09-26 02:25:28
Output files
Finished2024-09-26 02:25:38
Saved logsjustin-logs:262644.0-justin-prod-sched01.dune.hep.ac.uk.logs.tgz
List job events     Wrapper job log

Jobscript log (last 10,000 characters)

DEBUG:root:gfal.NoRename: checking if file exists davs://webdav.echo.stfc.ac.uk:1094/dune:/protodune/RSE/testpro/31/75/awt-1727313846-7X4qp65zlK
DEBUG:root:put: Attempt 1
DEBUG:root:gfal.NoRename: uploading file from awt-1727313846-7X4qp65zlK to davs://webdav.echo.stfc.ac.uk:1094/dune:/protodune/RSE/testpro/31/75/awt-1727313846-7X4qp65zlK
INFO:root:Successful upload of temporary file. davs://webdav.echo.stfc.ac.uk:1094/dune:/protodune/RSE/testpro/31/75/awt-1727313846-7X4qp65zlK
DEBUG:root:skip_upload_stat=False
DEBUG:root:stat: pfn=davs://webdav.echo.stfc.ac.uk:1094/dune:/protodune/RSE/testpro/31/75/awt-1727313846-7X4qp65zlK
DEBUG:root:gfal.NoRename: getting stats of file davs://webdav.echo.stfc.ac.uk:1094/dune:/protodune/RSE/testpro/31/75/awt-1727313846-7X4qp65zlK
DEBUG:root:Filesize: Expected=26 Found=26
DEBUG:root:Checksum: Expected=5e8706fb Found=5e8706fb
DEBUG:root:gfal.NoRename: closing protocol connection
DEBUG:root:Upload done.
INFO:root:Successfully uploaded file awt-1727313846-7X4qp65zlK
DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): dune-rucio.fnal.gov:443
/cvmfs/dune.opensciencegrid.org/products/dune/rucio/v34_4_2/NULL/lib/python3.9/site-packages/urllib3/connectionpool.py:1061: InsecureRequestWarning: Unverified HTTPS request is being made to host 'dune-rucio.fnal.gov'. Adding certificate verification is strongly advised. See: https://urllib3.readthedocs.io/en/1.26.x/advanced-usage.html#ssl-warnings
  warnings.warn(
DEBUG:urllib3.connectionpool:https://dune-rucio.fnal.gov:443 "POST /traces/ HTTP/1.1" 404 207
DEBUG:dogpile.lock:value creation lock <dogpile.cache.region.CacheRegion._LockWrapper object at 0x152ed81b8e20> acquired
DEBUG:dogpile.lock:Calling creation function for previously expired value
DEBUG:dogpile.cache.region:Cache value generated in 0.000 seconds for key(s): "host_to_choose_choice['https://dune-rucio.fnal.gov']"
DEBUG:dogpile.lock:Released creation lock
DEBUG:urllib3.connectionpool:Resetting dropped connection: dune-rucio.fnal.gov
DEBUG:urllib3.connectionpool:https://dune-rucio.fnal.gov:443 "PUT /replicas HTTP/1.1" 503 299
2024-09-25 19:22:58,555	WARNING	Waiting 0.25s due to reason: server returned 503 
WARNING:baseclient:Waiting 0.25s due to reason: server returned 503 
DEBUG:urllib3.connectionpool:Resetting dropped connection: dune-rucio.fnal.gov
DEBUG:urllib3.connectionpool:https://dune-rucio.fnal.gov:443 "PUT /replicas HTTP/1.1" 200 0
DEBUG:urllib3.connectionpool:https://dune-rucio.fnal.gov:443 "POST /dids/testpro/awt-uploads-202439/dids HTTP/1.1" 201 7
DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): dune-rucio.fnal.gov:443
DEBUG:urllib3.connectionpool:https://dune-rucio.fnal.gov:443 "POST /replicas/list HTTP/1.1" 503 299
2024-09-25 19:23:34,795	WARNING	Waiting 0.25s due to reason: server returned 503 
WARNING:baseclient:Waiting 0.25s due to reason: server returned 503 
DEBUG:urllib3.connectionpool:Starting new HTTPS connection (2): dune-rucio.fnal.gov:443
DEBUG:urllib3.connectionpool:https://dune-rucio.fnal.gov:443 "POST /replicas/list HTTP/1.1" 503 299
2024-09-25 19:23:50,461	WARNING	Waiting 0.5s due to reason: server returned 503 
WARNING:baseclient:Waiting 0.5s due to reason: server returned 503 
DEBUG:urllib3.connectionpool:Starting new HTTPS connection (3): dune-rucio.fnal.gov:443
DEBUG:urllib3.connectionpool:https://dune-rucio.fnal.gov:443 "POST /replicas/list HTTP/1.1" 503 299
2024-09-25 19:24:06,656	WARNING	Waiting 1.0s due to reason: server returned 503 
WARNING:baseclient:Waiting 1.0s due to reason: server returned 503 
--- Upload try 1/1
--- Rucio upload 1/1 returns 0
--- Replica check try 1/1
--- Rucio list_replicas call fails: An unknown exception occurred.
Details: no error information passed (http status code: 503)
--- No replica in Rucio, exit 98
'justin-rucio-upload --rse RAL_ECHO --protocol davs --scope testpro --dataset awt-uploads-202439 awt-1727313846-7X4qp65zlK --timeout 1200' returns 98


---------------------------------------------------------------------
US_UCSD SURFSARA davs root://penguin12.grid.surfsara.nl:21094/pnfs/grid.sara.nl/data/dune/disk/RSE/testpro/bb/7f/awt-download-2023-03-07-01.txt
'xrdcp --force --nopbar --verbose root://penguin12.grid.surfsara.nl:21094/pnfs/grid.sara.nl/data/dune/disk/RSE/testpro/bb/7f/awt-download-2023-03-07-01.txt downloaded.txt' returns 0
GFAL_CONFIG_DIR:    GFAL_PLUGIN_DIR: 
justin-rucio-upload attempt 1
DEBUG:root:Num. of files that upload client is processing: 1
DEBUG:dogpile.cache.region:No value present for key: "host_to_choose_choice['https://dune-rucio.fnal.gov']"
DEBUG:dogpile.lock:NeedRegenerationException
DEBUG:dogpile.lock:no value, waiting for create lock
DEBUG:dogpile.lock:value creation lock <dogpile.cache.region.CacheRegion._LockWrapper object at 0x14a5ebf478e0> acquired
DEBUG:dogpile.cache.region:No value present for key: "host_to_choose_choice['https://dune-rucio.fnal.gov']"
DEBUG:dogpile.lock:Calling creation function for not-yet-present value
DEBUG:dogpile.cache.region:Cache value generated in 0.000 seconds for key(s): "host_to_choose_choice['https://dune-rucio.fnal.gov']"
DEBUG:dogpile.lock:Released creation lock
DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): dune-rucio.fnal.gov:443
DEBUG:urllib3.connectionpool:https://dune-rucio.fnal.gov:443 "GET /rses/?expression=SURFSARA HTTP/1.1" 401 106
DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): dune-rucio.fnal.gov:443
DEBUG:urllib3.connectionpool:https://dune-rucio.fnal.gov:443 "GET /auth/x509_proxy HTTP/1.1" 503 299
2024-09-25 19:24:53,932	WARNING	Waiting 0.25s due to reason: server returned 503 
WARNING:baseclient:Waiting 0.25s due to reason: server returned 503 
DEBUG:urllib3.connectionpool:Starting new HTTPS connection (2): dune-rucio.fnal.gov:443
DEBUG:urllib3.connectionpool:https://dune-rucio.fnal.gov:443 "GET /auth/x509_proxy HTTP/1.1" 503 299
2024-09-25 19:25:09,962	WARNING	Waiting 0.5s due to reason: server returned 503 
WARNING:baseclient:Waiting 0.5s due to reason: server returned 503 
DEBUG:urllib3.connectionpool:Starting new HTTPS connection (3): dune-rucio.fnal.gov:443
DEBUG:urllib3.connectionpool:https://dune-rucio.fnal.gov:443 "GET /auth/x509_proxy HTTP/1.1" 503 299
2024-09-25 19:25:25,973	WARNING	Waiting 1.0s due to reason: server returned 503 
WARNING:baseclient:Waiting 1.0s due to reason: server returned 503 
--- Upload try 1/1
--- Rucio upload 1/1 fails: An unknown exception occurred.
Details: no error information passed (http status code: 503)
--- Exit with 99
'justin-rucio-upload --rse SURFSARA --protocol davs --scope testpro --dataset awt-uploads-202439 awt-1727313846-t7PZNmyN9L --timeout 1200' returns 99


subject   : /C=UK/O=eScience/OU=Manchester/L=HEP/CN=justin-jobs-production.dune.hep.ac.uk/CN=2021060471/CN=172731383895
issuer    : /C=UK/O=eScience/OU=Manchester/L=HEP/CN=justin-jobs-production.dune.hep.ac.uk/CN=2021060471
identity  : /C=UK/O=eScience/OU=Manchester/L=HEP/CN=justin-jobs-production.dune.hep.ac.uk/CN=2021060471
type      : RFC compliant proxy
strength  : 2048 bits
path      : /home/awt-proxy.pem
timeleft  : 166:58:31
key usage : Digital Signature, Key Encipherment, Key Agreement
=== VO dune extension information ===
VO        : dune
subject   : /C=UK/O=eScience/OU=Manchester/L=HEP/CN=justin-jobs-production.dune.hep.ac.uk
issuer    : /DC=org/DC=incommon/C=US/ST=Illinois/O=Fermi Research Alliance/CN=voms1.fnal.gov
attribute : /dune/Role=Production/Capability=NULL
attribute : /dune/Role=NULL/Capability=NULL
timeleft  : 145:11:36
uri       : voms1.fnal.gov:15042

===== Results =====

Download/upload commands:
xrdcp --force --nopbar --verbose $read_pfn downloaded.txt
justin-rucio-upload --rse $rse_name --protocol $write_protocol --scope testpro --dataset  --timeout 1200 FILENAME
Use the wrapper job link on the page for the job on the justIN Dashboard to find the full log file, with errors from these commands

Each line: $JUSTIN_SITE_NAME $rse_name $download_retval $upload_retval $read_pfn $write_protocol
==awt== US_UCSD DUNE_CERN_EOS 0 99 root://eospublic.cern.ch:1094//eos/experiment/neutplatform/protodune/dune/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
==awt== US_UCSD DUNE_ES_PIC 0 99 root://xrootd.pic.es:1094/pnfs/pic.es/data/dune/RSE/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
==awt== US_UCSD DUNE_FR_CCIN2P3_DISK 0 1 root://ccxrootdegee.in2p3.fr:1094/pnfs/in2p3.fr/data/dune/disk/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
==awt== US_UCSD DUNE_UK_LANCASTER_CEPH 0 98 root://xgate.hec.lancs.ac.uk:1094//cephfs/grid/dune/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
==awt== US_UCSD DUNE_US_BNL_SDCC 0 99 root://dcdndoor.sdcc.bnl.gov:1094//pnfs/sdcc.bnl.gov/data/dune/RSE/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
==awt== US_UCSD DUNE_US_FNAL_DISK_STAGE 0 99 root://fndca1.fnal.gov:1094/pnfs/fnal.gov/usr/dune/persistent/staging/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
==awt== US_UCSD MANCHESTER 54 99 root://bohr3226.tier2.hep.manchester.ac.uk:1094//dune/RSE/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
==awt== US_UCSD NIKHEF 0 99 root://dune.dcache.nikhef.nl:1094/pnfs/nikhef.nl/data/dune/generic/rucio/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
==awt== US_UCSD PRAGUE 0 0 root://golias100.farm.particle.cz:1094/dpm/farm.particle.cz/home/dune/RSE/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
==awt== US_UCSD QMUL 51 99 root://xrootd01.esc.qmul.ac.uk:1094//dune/RSE/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
==awt== US_UCSD RAL-PP 0 99 root://mover.pp.rl.ac.uk:1094/pnfs/pp.rl.ac.uk/data/dune/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
==awt== US_UCSD RAL_ECHO 0 98 root://xrootd.echo.stfc.ac.uk:1094/dune:/protodune/RSE/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
==awt== US_UCSD SURFSARA 0 99 root://penguin12.grid.surfsara.nl:21094/pnfs/grid.sara.nl/data/dune/disk/RSE/testpro/bb/7f/awt-download-2023-03-07-01.txt davs
justIN time: 2024-11-17 03:59:14 UTC       justIN version: 01.01.09