Difference between revisions of "Agenda-2020-09-14"
Jump to navigation
Jump to search
(46 intermediate revisions by 2 users not shown) | |||
Line 10: | Line 10: | ||
* plans on CentOS8 STARTED | * plans on CentOS8 STARTED | ||
** https://wiki.egi.eu/wiki/Next_middleware_release | ** https://wiki.egi.eu/wiki/Next_middleware_release | ||
* UMD-4.11.0 - June 29th, 2020 (regular release) | |||
** APEL-SSM 2.4.1 - Secure Stomp Messenger (SSM) updated - Added a delay when receiver is reconnecting to improve reliability. Improved the log output for SSM receivers so that there are fewer trivial entries and so it's more useful in tracking messages on the filesystem. | |||
argo-ams-library 0.5.1 - included in UMD | |||
** cvmfs 2.7.2 - CernVM File System (CernVM-FS) 2.7.2 is a patch release. It contains bugfixes and improvements for clients and servers. See https://cvmfs.readthedocs.io/en/2.7/cpt-releasenotes.html | |||
** CERN Frontier 4.11-3.1 - includes a patch and some fixes http://frontier.cern.ch/dist/rpms-debug/frontier-squidRELEASE_NOTES | |||
** FTS3 3.9.4 - includes bug fixes and improvements https://fts.web.cern.ch/sites/fts.web.cern.ch/themes/fts-webpage/releases-jekyll/releases/2020/05/07/FTS_3_9_4/ | |||
** dCache 5.2.20 - some fixes and improvements https://www.dcache.org/old/downloads/1.9/release-notes-5.2.shtml#20 | |||
* UMD-4.11.1 - July 21st, 2020 (emergency release) | |||
** This release contains includes an update of CERN Frontier for CentOS7 and SL6. | |||
* UMD-4.11.2 - Aug 14th, 2020 (emergency release) | |||
** This release contains includes an update of dCache for both CentOS7 and SL6. | |||
== Preview repository == | == Preview repository == | ||
*released on 2020-05 | * released on 2020-08-05 | ||
** '''[[Preview 1. | ** '''[[Preview 1.28.0]]''' [https://appdb.egi.eu/store/software/preview.repository/releases/1.0/1.28.0/ AppDB info] (sl6): dCache 5.2.25, frontier-squid 4.12.2, gfal2 2.18.1, xrootd 5.0.0 | ||
** '''[[Preview 2. | ** '''[[Preview 2.28.0]]''' [https://appdb.egi.eu/store/software/preview.repository/releases/2.0/2.28.0/ AppDB info] (CentOS 7): dCache 5.2.25, frontier-squid 4.12.2, gfal2 2.18.1, xrootd 5.0.0 | ||
= Operations = | = Operations = | ||
Line 20: | Line 34: | ||
== ARGO/SAM == | == ARGO/SAM == | ||
* [https://argo-mon-fedcloud.cro-ngi.hr/nagios/cgi-bin/status.cgi?servicegroup=SERVICE_org.opensciencegrid.htcondorce&style=detail HTCondor-CE probes] included in the ARGO_MON_OPERATORS profile on May 13th: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146949 | * [https://argo-mon-fedcloud.cro-ngi.hr/nagios/cgi-bin/status.cgi?servicegroup=SERVICE_org.opensciencegrid.htcondorce&style=detail HTCondor-CE probes] included in the ARGO_MON_OPERATORS profile on May 13th: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146949 | ||
** | ** '''(14th Sept)''' 70 endpoints, 14 CRITICAL, success rate is about 80% | ||
** '''on | ** '''on Oct 1st they will be included in the [https://poem.egi.eu/ui/public_metricprofiles/ARGO_MON_CRITICAL ARGO_MON_CRITICAL] profile (A/R computation)''' | ||
*** '''please fix the failures by that date''' | *** '''please fix the failures by that date''' | ||
** working on the probe for the host certificate validity check: [https://ggus.eu/index.php?mode=ticket_info&ticket_id=147386 GGUS 147386] | ** working on the probe for the host certificate validity check: [https://ggus.eu/index.php?mode=ticket_info&ticket_id=147386 GGUS 147386] | ||
* CREAM-CE metrics in the ARGO_MON_OPERATORS profile on [https://ggus.eu/index.php?mode=ticket_info&ticket_id=147169 May 27th]: eu.egi.CREAMCE-JobSubmit, eu.egi.CREAMCE.WN-Csh, eu.egi.CREAMCE.WN-Softver | * CREAM-CE metrics in the ARGO_MON_OPERATORS profile on [https://ggus.eu/index.php?mode=ticket_info&ticket_id=147169 May 27th]: eu.egi.CREAMCE-JobSubmit, eu.egi.CREAMCE.WN-Csh, eu.egi.CREAMCE.WN-Softver | ||
** [https://argo-mon.egi.eu/nagios/cgi-bin/status.cgi?servicegroup=SERVICE_CREAM-CE&style=detail results]: | ** '''(14th Sept)''' [https://argo-mon.egi.eu/nagios/cgi-bin/status.cgi?servicegroup=SERVICE_CREAM-CE&style=detail results]: 152 endpoints, 20 WARNING (Timeout occurred (900 sec) ), 24 CRITICAL. Success rate 84.2% (71.1% including the WARNING) | ||
*When eu.egi.CREAMCE.WN-Softver is successful: | *When eu.egi.CREAMCE.WN-Softver is successful: | ||
CREAM JobOutput OK: retrieved outputSandbox: ['std.err', 'std.out'] | CREAM JobOutput OK: retrieved outputSandbox: ['std.err', 'std.out'] | ||
Line 69: | Line 83: | ||
[https://wiki.egi.eu/wiki/Verify_Configuration_Records#2020-05 List of tickets]. | [https://wiki.egi.eu/wiki/Verify_Configuration_Records#2020-05 List of tickets]. | ||
* 30 tickets | * 30 tickets | ||
* Not yet solved | * Not yet solved: 8 | ||
== Monthly Availability/Reliability == | == Monthly Availability/Reliability == | ||
Line 75: | Line 89: | ||
*Under-performed sites in the past A/R reports with issues not yet fixed: | *Under-performed sites in the past A/R reports with issues not yet fixed: | ||
** AsiaPacific: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147748 | ** AsiaPacific: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147748 | ||
*** '''HK-HKU-CC-01''' | *** '''HK-HKU-CC-01''': migrating DPM from sl6 to CenOS7 | ||
*** '''TW-NCUHEP''' | *** '''TW-NCUHEP''': ARC-CE failures due to outdated CAs package | ||
** NGI_BG: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147747 | ** NGI_BG: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147747 | ||
*** '''BG01-IPP''' | *** '''BG01-IPP''': CREAM-CE failures, improving... | ||
**NGI_DE: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146871 | **NGI_DE: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146871 | ||
***'''GoeGRID''': CREAM-CE intermittent failures not affecting ATLAS; failures with ARC-CE | ***'''GoeGRID''': CREAM-CE intermittent failures not affecting ATLAS; failures with ARC-CE | ||
Line 84: | Line 98: | ||
***'''mainz''': some problems in March and April, that could not be fixed easily; in May, the HPC infrastructure was attacked and the whole computer center was shut down; in downtime. | ***'''mainz''': some problems in March and April, that could not be fixed easily; in May, the HPC infrastructure was attacked and the whole computer center was shut down; in downtime. | ||
***'''wuppertalprod''': SRM failures to to a BDII issue, fixed | ***'''wuppertalprod''': SRM failures to to a BDII issue, fixed | ||
** NGI_GRNET: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148171 | |||
*** HG-02-IASA: problems with certificates renewal due to COVID situation; the montlhy figures are improving | |||
**NGI_IT: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148170 | |||
***Hephy-Vienna: SRM decommissioned, moved to EOS | |||
***INFN-PADOVA-STACK: the ESACO instance was having problems with the new AAI host certificate. See [https://ggus.eu/index.php?mode=ticket_info&ticket_id=148242 148242] | |||
**NGI_PL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147311 | **NGI_PL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147311 | ||
***WCSS64 | ***WCSS64 | ||
** NGI_PL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148167 | |||
*** WUT: downtime for site update, production jobs can run. | |||
**NGI_UK: | **NGI_UK: | ||
***'''UKI-NORTHGRID-SHEF-HEP''': https://ggus.eu/index.php?mode=ticket_info&ticket_id=146455 ARC-CE re-installed, some condor problems to fix | ***'''UKI-NORTHGRID-SHEF-HEP''': https://ggus.eu/index.php?mode=ticket_info&ticket_id=146455 ARC-CE re-installed, some condor problems to fix | ||
Line 91: | Line 112: | ||
** NGI_UA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147750 | ** NGI_UA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147750 | ||
*** '''UA-ISMA''': migration to ARC6 and other planned software updates | *** '''UA-ISMA''': migration to ARC6 and other planned software updates | ||
*Under-performed sites after 3 consecutive months, under-performed NGIs, QoS violations: (''' | *Under-performed sites after 3 consecutive months, under-performed NGIs, QoS violations: ('''August 2020'''): | ||
** | ** CERN: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148516 | ||
*** | *** webdav failures due to insufficient space in the partition, fixed. | ||
** NGI_DE: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148519 | |||
*** LRZ-LMU: CE had problems due to the decommission of SharedFS | |||
** NGI_HR: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148518 | |||
*** egee.irb.hr: in the process of a major upgrade from CentOS 6 to CentOS 7, some delays. | |||
** NGI_IL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148521 | |||
*** TECHNION-HEP | |||
** NGI_NL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148520 | |||
*** SARA-MATRIX | |||
** ROC_LA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148515 | |||
*** ATLAND: downtime due to powercut and quarantine | |||
*sites suspended: | *sites suspended: | ||
Line 100: | Line 132: | ||
* please provide updates to the IPv6 assessment (ongoing) https://wiki.egi.eu/w/index.php?title=IPV6_Assessment | * please provide updates to the IPv6 assessment (ongoing) https://wiki.egi.eu/w/index.php?title=IPV6_Assessment | ||
* if any relevant, information will be summarised at OMB | * if any relevant, information will be summarised at OMB | ||
== CREAM-CE Decommission == | |||
* End of Security Updates and Support: 31st Dec 2020 (Decommissioning deadline) | |||
** Original broadcast: https://operations-portal.egi.eu/broadcast/archive/2293 | |||
* [https://wiki.egi.eu/wiki/PROC16_Decommissioning_of_unsupported_software PROC16 Decommission of unsupported software] | |||
* Decommissioning start date: Oct 1st 2020 | |||
** a probe detecting CREAM-CE endpoints will be run, returning WARNING status | |||
* Nov 1st: probe returns CRITICAL status, alarms created on the ROD dashboard, ROD teams start to create tickets | |||
* 1st Jan 2021: EGI Ops will start chasing the sites still providing CREAM-CE endpoints | |||
** By this time service end-points which couldn't be upgraded should be put into downtime by site admin or ROD: | |||
== ARC Middleware 5 end of support, migration to ARC 6 == | == ARC Middleware 5 end of support, migration to ARC 6 == | ||
Line 115: | Line 158: | ||
|- | |- | ||
| 2020-07-13 || 53 || 29 || - | | 2020-07-13 || 53 || 29 || - | ||
|- | |||
| 2020-09-14 || 34 || 18 || - | |||
|} | |} | ||
== Storage accounting == | == Storage accounting == | ||
Many | Many sites stopped the publication of storage accounting records. Opened [https://ggus.eu/index.php?mode=ticket_search&show_columns_check%5B0%5D=TICKET_TYPE&show_columns_check%5B1%5D=AFFECTED_VO&show_columns_check%5B2%5D=AFFECTED_SITE&show_columns_check%5B3%5D=PRIORITY&show_columns_check%5B4%5D=RESPONSIBLE_UNIT&show_columns_check%5B5%5D=STATUS&show_columns_check%5B6%5D=DATE_OF_CHANGE&show_columns_check%5B7%5D=SHORT_DESCRIPTION&show_columns_check%5B8%5D=SCOPE&su_hierarchy=0&keyword=publishing+storage+accounting+records&specattrib=none&status=all&typeofproblem=all&ticket_category=all&date_type=creation+date&tf_radio=1&timeframe=any&from_date=10+Jul+2020&to_date=11+Jul+2020&orderticketsby=REQUEST_ID&orderhow=desc&search_submit=GO%21&ticket_per_page=60 57 tickets] to fix that. | ||
* 15 tickets not solved yet | |||
* page for checking when the records were published: http://goc-accounting.grid-support.ac.uk/storagetest/storagesitesystems.html | * page for checking when the records were published: http://goc-accounting.grid-support.ac.uk/storagetest/storagesitesystems.html | ||
* [http://accounting-devel.egi.eu/storage.php Accounting Portal Prototype view] | * [http://accounting-devel.egi.eu/storage.php Accounting Portal Prototype view] | ||
= AOB = | = AOB = | ||
Line 133: | Line 173: | ||
== Next meeting == | == Next meeting == | ||
Oct 12th, 2020 https://indico.egi.eu/event/5099/ |
Latest revision as of 11:33, 22 September 2020
Main | EGI.eu operations services | Support | Documentation | Tools | Activities | Performance | Technology | Catch-all Services | Resource Allocation | Security |
Documentation menu: | Home • | Manuals • | Procedures • | Training • | Other • | Contact ► | For: | VO managers • | Administrators |
Back to https://wiki.egi.eu/wiki/Operations_Meeting
General information
Middleware
UMD
- plans on CentOS8 STARTED
- UMD-4.11.0 - June 29th, 2020 (regular release)
- APEL-SSM 2.4.1 - Secure Stomp Messenger (SSM) updated - Added a delay when receiver is reconnecting to improve reliability. Improved the log output for SSM receivers so that there are fewer trivial entries and so it's more useful in tracking messages on the filesystem.
argo-ams-library 0.5.1 - included in UMD
- cvmfs 2.7.2 - CernVM File System (CernVM-FS) 2.7.2 is a patch release. It contains bugfixes and improvements for clients and servers. See https://cvmfs.readthedocs.io/en/2.7/cpt-releasenotes.html
- CERN Frontier 4.11-3.1 - includes a patch and some fixes http://frontier.cern.ch/dist/rpms-debug/frontier-squidRELEASE_NOTES
- FTS3 3.9.4 - includes bug fixes and improvements https://fts.web.cern.ch/sites/fts.web.cern.ch/themes/fts-webpage/releases-jekyll/releases/2020/05/07/FTS_3_9_4/
- dCache 5.2.20 - some fixes and improvements https://www.dcache.org/old/downloads/1.9/release-notes-5.2.shtml#20
- UMD-4.11.1 - July 21st, 2020 (emergency release)
- This release contains includes an update of CERN Frontier for CentOS7 and SL6.
- UMD-4.11.2 - Aug 14th, 2020 (emergency release)
- This release contains includes an update of dCache for both CentOS7 and SL6.
Preview repository
- released on 2020-08-05
- Preview 1.28.0 AppDB info (sl6): dCache 5.2.25, frontier-squid 4.12.2, gfal2 2.18.1, xrootd 5.0.0
- Preview 2.28.0 AppDB info (CentOS 7): dCache 5.2.25, frontier-squid 4.12.2, gfal2 2.18.1, xrootd 5.0.0
Operations
ARGO/SAM
- HTCondor-CE probes included in the ARGO_MON_OPERATORS profile on May 13th: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146949
- (14th Sept) 70 endpoints, 14 CRITICAL, success rate is about 80%
- on Oct 1st they will be included in the ARGO_MON_CRITICAL profile (A/R computation)
- please fix the failures by that date
- working on the probe for the host certificate validity check: GGUS 147386
- CREAM-CE metrics in the ARGO_MON_OPERATORS profile on May 27th: eu.egi.CREAMCE-JobSubmit, eu.egi.CREAMCE.WN-Csh, eu.egi.CREAMCE.WN-Softver
- (14th Sept) results: 152 endpoints, 20 WARNING (Timeout occurred (900 sec) ), 24 CRITICAL. Success rate 84.2% (71.1% including the WARNING)
- When eu.egi.CREAMCE.WN-Softver is successful:
CREAM JobOutput OK: retrieved outputSandbox: ['std.err', 'std.out'] **** std.err **** **** std.out **** egee01 has UMD 3.14.4
When it fails:
CREAM JobOutput ERROR [DONE-OK, exitCode=1 ]: retrieved outputSandbox: ['std.err', 'std.out'] **** std.err **** **** std.out **** ERROR: unable to find glite, EMI, LCG or UMD WN version on n1037-amd
FedCloud
Feedback from DMSU
Verify configuration records
On a yearly basis, the information registered into GOC-DB need to be verified. NGIs and RCs have been asked to check them. In particular:
- NGI managers should review the people registered and the roles assigned to them, and in particular check the following information:
- ROD E-Mail
- Security E-Mail
- NGI Managers should also review the status of the "not certified" RCs, in according to the RC Status Workflow;
- RCs administrators should review the people registered and the roles assigned to them, and in particular check the following information:
- telephone numbers
- CSIRT E-Mail
- RC administrators should also review the information related to the registered service endpoints.
The process should be completed by June 22nd.
- 30 tickets
- Not yet solved: 8
Monthly Availability/Reliability
- Under-performed sites in the past A/R reports with issues not yet fixed:
- AsiaPacific: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147748
- HK-HKU-CC-01: migrating DPM from sl6 to CenOS7
- TW-NCUHEP: ARC-CE failures due to outdated CAs package
- NGI_BG: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147747
- BG01-IPP: CREAM-CE failures, improving...
- NGI_DE: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146871
- GoeGRID: CREAM-CE intermittent failures not affecting ATLAS; failures with ARC-CE
- NGI_DE: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147313
- mainz: some problems in March and April, that could not be fixed easily; in May, the HPC infrastructure was attacked and the whole computer center was shut down; in downtime.
- wuppertalprod: SRM failures to to a BDII issue, fixed
- NGI_GRNET: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148171
- HG-02-IASA: problems with certificates renewal due to COVID situation; the montlhy figures are improving
- NGI_IT: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148170
- Hephy-Vienna: SRM decommissioned, moved to EOS
- INFN-PADOVA-STACK: the ESACO instance was having problems with the new AAI host certificate. See 148242
- NGI_PL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147311
- WCSS64
- NGI_PL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148167
- WUT: downtime for site update, production jobs can run.
- NGI_UK:
- UKI-NORTHGRID-SHEF-HEP: https://ggus.eu/index.php?mode=ticket_info&ticket_id=146455 ARC-CE re-installed, some condor problems to fix
- UKI-SOUTHGRID-SUSX: https://ggus.eu/index.php?mode=ticket_info&ticket_id=144720 Migration from CREAM to ARC, WN migration to CentOS7; SRM to be decommissioned; ARC-CE was failing the IGTF test, then solved; site-bdii failures.
- NGI_UA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147750
- UA-ISMA: migration to ARC6 and other planned software updates
- AsiaPacific: https://ggus.eu/index.php?mode=ticket_info&ticket_id=147748
- Under-performed sites after 3 consecutive months, under-performed NGIs, QoS violations: (August 2020):
- CERN: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148516
- webdav failures due to insufficient space in the partition, fixed.
- NGI_DE: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148519
- LRZ-LMU: CE had problems due to the decommission of SharedFS
- NGI_HR: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148518
- egee.irb.hr: in the process of a major upgrade from CentOS 6 to CentOS 7, some delays.
- NGI_IL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148521
- TECHNION-HEP
- NGI_NL: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148520
- SARA-MATRIX
- ROC_LA: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148515
- ATLAND: downtime due to powercut and quarantine
- CERN: https://ggus.eu/index.php?mode=ticket_info&ticket_id=148516
- sites suspended:
IPv6 readiness plans
- please provide updates to the IPv6 assessment (ongoing) https://wiki.egi.eu/w/index.php?title=IPV6_Assessment
- if any relevant, information will be summarised at OMB
CREAM-CE Decommission
- End of Security Updates and Support: 31st Dec 2020 (Decommissioning deadline)
- Original broadcast: https://operations-portal.egi.eu/broadcast/archive/2293
- PROC16 Decommission of unsupported software
- Decommissioning start date: Oct 1st 2020
- a probe detecting CREAM-CE endpoints will be run, returning WARNING status
- Nov 1st: probe returns CRITICAL status, alarms created on the ROD dashboard, ROD teams start to create tickets
- 1st Jan 2021: EGI Ops will start chasing the sites still providing CREAM-CE endpoints
- By this time service end-points which couldn't be upgraded should be put into downtime by site admin or ROD:
ARC Middleware 5 end of support, migration to ARC 6
- EGI Operations Broadcast
- PROC16 Decommission of unsupported software
- deadline: end of July
- Catalin is in contact with ARC team to get a webinar on ARC administration, scheduled (to be confirmed) for July 6th please contact operations@ for information
- Status
Date | Number of endpoints in BDII | Number of GGUS tickets | Issues |
---|---|---|---|
2020-06-08 | 75 | 42 | Some ARC endpoints publish a timestamp instead of a version like 5.X.Y; we can fairly assume they are ARC6 nightly builds, but we're going to close the corresponding tickets after explicit confirmation from the site admin. |
2020-07-13 | 53 | 29 | - |
2020-09-14 | 34 | 18 | - |
Storage accounting
Many sites stopped the publication of storage accounting records. Opened 57 tickets to fix that.
- 15 tickets not solved yet
- page for checking when the records were published: http://goc-accounting.grid-support.ac.uk/storagetest/storagesitesystems.html
- Accounting Portal Prototype view
AOB
Next meeting
Oct 12th, 2020 https://indico.egi.eu/event/5099/