Alert.png The wiki is deprecated and due to be decommissioned by the end of September 2022.
The content is being migrated to other supports, new updates will be ignored and lost.
If needed you can get in touch with EGI SDIS team using operations @ egi.eu.

Difference between revisions of "Agenda-2020-10-12"

From EGIWiki
Jump to navigation Jump to search
 
(18 intermediate revisions by 2 users not shown)
Line 8: Line 8:


== UMD ==
== UMD ==
* plans on CentOS8 STARTED
* plans on CentOS8 ONGOING
** https://wiki.egi.eu/wiki/Next_middleware_release
** https://wiki.egi.eu/wiki/Next_middleware_release


* UMD-4.11.0 - June 29th, 2020 (regular release)
* UMD-4.12.0 regular release is almost ready (testing RC)
** APEL-SSM 2.4.1 - Secure Stomp Messenger (SSM) updated - Added a delay when receiver is reconnecting to improve reliability. Improved the log output for SSM receivers so that there are fewer trivial entries and so it's more useful in tracking messages on the filesystem.
** CVMFS 2.7.3, ARCCE 6.7.0, gfal 2.18.1, davix 0.7.6, xrootd 4.12.3
argo-ams-library 0.5.1 - included in UMD
** next releases: update for VOMS on C7, StoRM on C7, BDII C7/SL6
** cvmfs 2.7.2 - CernVM File System (CernVM-FS) 2.7.2 is a patch release. It contains bugfixes and improvements for clients and servers. See https://cvmfs.readthedocs.io/en/2.7/cpt-releasenotes.html
** CERN Frontier 4.11-3.1 - includes a patch and some fixes http://frontier.cern.ch/dist/rpms-debug/frontier-squidRELEASE_NOTES
** FTS3 3.9.4 - includes bug fixes and improvements https://fts.web.cern.ch/sites/fts.web.cern.ch/themes/fts-webpage/releases-jekyll/releases/2020/05/07/FTS_3_9_4/
** dCache 5.2.20 - some fixes and improvements https://www.dcache.org/old/downloads/1.9/release-notes-5.2.shtml#20


* UMD-4.11.1 - July 21st, 2020 (emergency release)
** This release contains includes an update of CERN Frontier for CentOS7 and SL6.


* UMD-4.11.2 - Aug 14th, 2020 (emergency release)
** This release contains includes an update of dCache for both CentOS7 and SL6.


== Preview repository  ==
== Preview repository  ==
* released on 2020-08-05
* released on 2020-10-09
** '''[[Preview 1.28.0]]''' [https://appdb.egi.eu/store/software/preview.repository/releases/1.0/1.28.0/ AppDB info] (sl6):  dCache 5.2.25, frontier-squid 4.12.2, gfal2 2.18.1, xrootd 5.0.0
** '''[[Preview 1.29.0]]''' [https://appdb.egi.eu/store/software/preview.repository/releases/1.0/1.29.0/ AppDB info] (sl6):  ARC 6.8.0 and 6.8.1, BDII 5.5.26, CVMFS 2.7.4, dCache 5.2.31, DMLite/DPM 1.14.0, frontier-squid 4.13.1, glite-info-update-endpoints 3.0.2, lcg-info 1.12.5, STORM 1.11.18
** '''[[Preview 2.28.0]]''' [https://appdb.egi.eu/store/software/preview.repository/releases/2.0/2.28.0/ AppDB info] (CentOS 7):  dCache 5.2.25, frontier-squid 4.12.2, gfal2 2.18.1, xrootd 5.0.0
** '''[[Preview 2.29.0]]''' [https://appdb.egi.eu/store/software/preview.repository/releases/2.0/2.29.0/ AppDB info] (CentOS 7):  ARC 6.8.0 and 6.8.1, BDII 5.5.26, CVMFS 2.7.4, dCache 5.2.31, DMLite/DPM 1.14.0, frontier-squid 4.13.1, glite-info-update-endpoints 3.0.2, lcg-info 1.12.5, STORM 1.11.18


= Operations  =
= Operations  =
Line 35: Line 27:
* [https://argo-mon-fedcloud.cro-ngi.hr/nagios/cgi-bin/status.cgi?servicegroup=SERVICE_org.opensciencegrid.htcondorce&style=detail HTCondor-CE probes] included in the ARGO_MON_OPERATORS profile on May 13th: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146949
* [https://argo-mon-fedcloud.cro-ngi.hr/nagios/cgi-bin/status.cgi?servicegroup=SERVICE_org.opensciencegrid.htcondorce&style=detail HTCondor-CE probes] included in the ARGO_MON_OPERATORS profile on May 13th: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146949
** '''(14th Sept)''' 70 endpoints, 14 CRITICAL, success rate is about 80%
** '''(14th Sept)''' 70 endpoints, 14 CRITICAL, success rate is about 80%
** '''on Oct 1st they will be included in the [https://poem.egi.eu/ui/public_metricprofiles/ARGO_MON_CRITICAL ARGO_MON_CRITICAL] profile (A/R computation)'''
** '''Oct 1st: included in the [https://poem.egi.eu/ui/public_metricprofiles/ARGO_MON_CRITICAL ARGO_MON_CRITICAL] profile (A/R computation)'''
*** '''please fix the failures by that date'''
*** (Oct 12th) 71 endpoints, success rate (including WARNING) 85.9%
** working on the probe for the host certificate validity check: [https://ggus.eu/index.php?mode=ticket_info&ticket_id=147386 GGUS 147386]
** working on the probe for the host certificate validity check: [https://ggus.eu/index.php?mode=ticket_info&ticket_id=147386 GGUS 147386]
* CREAM-CE metrics in the ARGO_MON_OPERATORS profile on [https://ggus.eu/index.php?mode=ticket_info&ticket_id=147169 May 27th]: eu.egi.CREAMCE-JobSubmit, eu.egi.CREAMCE.WN-Csh, eu.egi.CREAMCE.WN-Softver
** '''(14th Sept)''' [https://argo-mon.egi.eu/nagios/cgi-bin/status.cgi?servicegroup=SERVICE_CREAM-CE&style=detail results]:  152 endpoints, 20 WARNING (Timeout occurred (900 sec) ), 24 CRITICAL. Success rate 84.2% (71.1% including the WARNING)
*When eu.egi.CREAMCE.WN-Softver is successful:
CREAM JobOutput OK: retrieved outputSandbox: ['std.err', 'std.out']
**** std.err ****
**** std.out ****
egee01 has UMD 3.14.4
When it fails:
CREAM JobOutput ERROR [DONE-OK, exitCode=1 ]: retrieved outputSandbox: ['std.err', 'std.out']
**** std.err ****
 
**** std.out ****
ERROR: unable to find glite, EMI, LCG or UMD WN version on n1037-amd


== FedCloud  ==
== FedCloud  ==
Line 63: Line 37:


== Feedback from DMSU  ==
== Feedback from DMSU  ==
== Verify configuration records ==
On a yearly basis, the information registered into GOC-DB need to be verified.
NGIs and RCs have been asked to check them. In particular:
# '''NGI managers should review the people registered and the roles assigned to them, and in particular check the following information:'''
#* E-Mail
#* ROD E-Mail
#* Security E-Mail
:NGI Managers should also review the status of the "not certified" RCs, in according to the [https://wiki.egi.eu/wiki/PROC09#Resource_Center_status_Workflow RC Status Workflow];
# '''RCs administrators should review the people registered and the roles assigned to them, and in particular check the following information:'''
#* E-Mail
#* telephone numbers
#* CSIRT E-Mail
: RC administrators should also review the information related to the registered service endpoints.
'''The process should be completed by June 22nd.'''
[https://wiki.egi.eu/wiki/Verify_Configuration_Records#2020-05 List of tickets].
* 30 tickets
* Not yet solved: 8


== Monthly Availability/Reliability ==
== Monthly Availability/Reliability ==
Line 91: Line 44:
*** '''HK-HKU-CC-01''': migrating DPM from sl6 to CenOS7
*** '''HK-HKU-CC-01''': migrating DPM from sl6 to CenOS7
*** '''TW-NCUHEP''': ARC-CE failures due to outdated CAs package
*** '''TW-NCUHEP''': ARC-CE failures due to outdated CAs package
** NGI_BG: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147747
*** '''BG01-IPP''': CREAM-CE failures, improving...
**NGI_DE: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146871
**NGI_DE: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146871
***'''GoeGRID''': CREAM-CE intermittent failures not affecting ATLAS; failures with ARC-CE
***'''GoeGRID''': CREAM-CE intermittent failures not affecting ATLAS; failures with ARC-CE, now passing the tests
** NGI_DE: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148519
** NGI_DE: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148519
*** LRZ-LMU: CE had problems due to the decommission of SharedFS; the other CE returns UNKNOWN in the IGTF test.
*** LRZ-LMU: CE had problems due to the decommission of SharedFS; the other CE returns UNKNOWN in the IGTF test.
Line 101: Line 52:
**NGI_PL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147311
**NGI_PL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147311
***WCSS64
***WCSS64
** NGI_PL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148167
*** WUT: downtime for site update, production jobs can run.
**NGI_UK:
**NGI_UK:
***'''UKI-NORTHGRID-SHEF-HEP''': https://ggus.eu/index.php?mode=ticket_info&ticket_id=146455 ARC-CE re-installed, some condor problems to fix  
***'''UKI-NORTHGRID-SHEF-HEP''': https://ggus.eu/index.php?mode=ticket_info&ticket_id=146455 ARC-CE re-installed, some condor problems to fix  
***'''UKI-SOUTHGRID-SUSX''': https://ggus.eu/index.php?mode=ticket_info&ticket_id=144720 Migration from CREAM to ARC, WN migration to CentOS7; SRM to be decommissioned; ARC-CE was failing the IGTF test, then solved; site-bdii failures.
***'''UKI-SOUTHGRID-SUSX''': https://ggus.eu/index.php?mode=ticket_info&ticket_id=144720 Migration from CREAM to ARC, WN migration to CentOS7; SRM to be decommissioned; ARC-CE was failing the IGTF test, then solved; site-bdii failures.
*Under-performed sites after 3 consecutive months, under-performed NGIs, QoS violations: ('''August 2020'''):
** NGI_NL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148520
*** SARA-MATRIX
** ROC_LA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148515
** ROC_LA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148515
*** ATLAND: downtime due to powercut and quarantine
*** ATLAND: downtime due to powercut and quarantine
*Under-performed sites after 3 consecutive months, under-performed NGIs, QoS violations: ('''September 2020'''):
** NGI_IT: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148957
*** INFN-CATANIA
*** INFN-CATANIA-STACK
*** INFN-PADOVA
** NGI_UA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148958
*** UA-NSCMBR: IGTF outdated
** ROC_LA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148956
*** CBPF


*sites suspended:
*sites suspended:
Line 127: Line 82:
* Decommissioning start date: Oct 1st 2020
* Decommissioning start date: Oct 1st 2020
** a probe detecting CREAM-CE endpoints will be run, returning WARNING status
** a probe detecting CREAM-CE endpoints will be run, returning WARNING status
** GGUS ticket: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148715
** [https://argo-mon.egi.eu/nagios/cgi-bin/status.cgi?servicegroup=SERVICE_CREAM-CE&style=detail eu.egi.sec.CREAMCE]
* Nov 1st: probe returns CRITICAL status, alarms created on the ROD dashboard, ROD teams start to create tickets
* Nov 1st: probe returns CRITICAL status, alarms created on the ROD dashboard, ROD teams start to create tickets
* 1st Jan 2021: EGI Ops will start chasing the sites still providing CREAM-CE endpoints
* 1st Jan 2021: EGI Ops will start chasing the sites still providing CREAM-CE endpoints
Line 147: Line 104:
|-
|-
| 2020-09-14 || 34 || 18 || -
| 2020-09-14 || 34 || 18 || -
|-
| 2020-10-12 || 32 || 19 || -
|}
|}


Line 152: Line 111:


Many sites stopped the publication of storage accounting records. Opened [https://ggus.eu/index.php?mode=ticket_search&show_columns_check%5B0%5D=TICKET_TYPE&show_columns_check%5B1%5D=AFFECTED_VO&show_columns_check%5B2%5D=AFFECTED_SITE&show_columns_check%5B3%5D=PRIORITY&show_columns_check%5B4%5D=RESPONSIBLE_UNIT&show_columns_check%5B5%5D=STATUS&show_columns_check%5B6%5D=DATE_OF_CHANGE&show_columns_check%5B7%5D=SHORT_DESCRIPTION&show_columns_check%5B8%5D=SCOPE&su_hierarchy=0&keyword=publishing+storage+accounting+records&specattrib=none&status=all&typeofproblem=all&ticket_category=all&date_type=creation+date&tf_radio=1&timeframe=any&from_date=10+Jul+2020&to_date=11+Jul+2020&orderticketsby=REQUEST_ID&orderhow=desc&search_submit=GO%21&ticket_per_page=60  57 tickets] to fix that.
Many sites stopped the publication of storage accounting records. Opened [https://ggus.eu/index.php?mode=ticket_search&show_columns_check%5B0%5D=TICKET_TYPE&show_columns_check%5B1%5D=AFFECTED_VO&show_columns_check%5B2%5D=AFFECTED_SITE&show_columns_check%5B3%5D=PRIORITY&show_columns_check%5B4%5D=RESPONSIBLE_UNIT&show_columns_check%5B5%5D=STATUS&show_columns_check%5B6%5D=DATE_OF_CHANGE&show_columns_check%5B7%5D=SHORT_DESCRIPTION&show_columns_check%5B8%5D=SCOPE&su_hierarchy=0&keyword=publishing+storage+accounting+records&specattrib=none&status=all&typeofproblem=all&ticket_category=all&date_type=creation+date&tf_radio=1&timeframe=any&from_date=10+Jul+2020&to_date=11+Jul+2020&orderticketsby=REQUEST_ID&orderhow=desc&search_submit=GO%21&ticket_per_page=60  57 tickets] to fix that.
* 15 tickets not solved yet
* 12 tickets not solved yet
* page for checking when the records were published: http://goc-accounting.grid-support.ac.uk/storagetest/storagesitesystems.html
* page for checking when the records were published: http://goc-accounting.grid-support.ac.uk/storagetest/storagesitesystems.html
* [http://accounting-devel.egi.eu/storage.php Accounting Portal Prototype view]
* [http://accounting-devel.egi.eu/storage.php Accounting Portal Prototype view]
Line 160: Line 119:


== Next meeting  ==
== Next meeting  ==
Oct 12th, 2020 https://indico.egi.eu/event/5099/
Nov 16th, 2020 https://indico.egi.eu/event/5100/

Latest revision as of 14:14, 12 October 2020

Main EGI.eu operations services Support Documentation Tools Activities Performance Technology Catch-all Services Resource Allocation Security


Documentation menu: Home Manuals Procedures Training Other Contact For: VO managers Administrators


Back to https://wiki.egi.eu/wiki/Operations_Meeting

General information

Middleware

UMD

  • UMD-4.12.0 regular release is almost ready (testing RC)
    • CVMFS 2.7.3, ARCCE 6.7.0, gfal 2.18.1, davix 0.7.6, xrootd 4.12.3
    • next releases: update for VOMS on C7, StoRM on C7, BDII C7/SL6


Preview repository

  • released on 2020-10-09
    • Preview 1.29.0 AppDB info (sl6): ARC 6.8.0 and 6.8.1, BDII 5.5.26, CVMFS 2.7.4, dCache 5.2.31, DMLite/DPM 1.14.0, frontier-squid 4.13.1, glite-info-update-endpoints 3.0.2, lcg-info 1.12.5, STORM 1.11.18
    • Preview 2.29.0 AppDB info (CentOS 7): ARC 6.8.0 and 6.8.1, BDII 5.5.26, CVMFS 2.7.4, dCache 5.2.31, DMLite/DPM 1.14.0, frontier-squid 4.13.1, glite-info-update-endpoints 3.0.2, lcg-info 1.12.5, STORM 1.11.18

Operations

ARGO/SAM

FedCloud

Feedback from DMSU

Monthly Availability/Reliability


  • sites suspended:

IPv6 readiness plans

CREAM-CE Decommission

ARC Middleware 5 end of support, migration to ARC 6

  • Status
Date Number of endpoints in BDII Number of GGUS tickets Issues
2020-06-08 75 42 Some ARC endpoints publish a timestamp instead of a version like 5.X.Y; we can fairly assume they are ARC6 nightly builds, but we're going to close the corresponding tickets after explicit confirmation from the site admin.
2020-07-13 53 29 -
2020-09-14 34 18 -
2020-10-12 32 19 -

Storage accounting

Many sites stopped the publication of storage accounting records. Opened 57 tickets to fix that.

AOB

Next meeting

Nov 16th, 2020 https://indico.egi.eu/event/5100/