No failure patterns match your search.
Environment: DEV
Failure patterns: 49 · Failures matched: 183 · Total runs: 103 · Runs affected: 83
| Failure patternThe canonical description of a recurring CI failure, extracted and normalized from raw logs. | Failed atThe stage of the job run where this failure occurred: 'provision' (environment setup, DEV only), 'e2e' (test suite execution), or 'other' (CI infrastructure issues that did not produce a failure pattern). | Number of distinct job runs where this failure pattern was detected. | Percentage of all job runs in this environment affected by this failure pattern during the selected window. | Classification of this failure pattern: Regression (likely caused by a specific PR), Flake (intermittent failure spread across days), Noise (low-quality or generic pattern), or Indeterminate. | TrendShows daily activity for this failure pattern in a trailing window anchored to the selected end date. The sparkline covers at least 7 days and at most 14 days, depending on the current window size. | Also inOther environments where the same failure pattern was also detected during the selected window. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Show signal detailsLikely regression — post-good=0; only seen in DEV; only seen in one PRERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable; host...full failure pattern: ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable; hosted cluster degraded: UnavailableReplicas: [router]; provider Microsoft.RedHatOpenShift Signal: RegressionSignal: Regression — post-good=0; only seen in DEV; only seen in one PR | e2e | 1 | 0.97%1 of 103 job runs affected | RegressionSignal: Regression — post-good=0; only seen in DEV; only seen in one PR | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 2 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (2)fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_feature_aggregation.go:184]: failed to create HCP cluster "agg-cluster" with aggregated settings
Unexpected error:
<*fmt.wrapError | 0xc00133a200>:
failed waiting for cluster="agg-cluster" in resourcegroup="feature-aggregation-qp6jkgrm7jtk" to finish creating: GET https://rp.j8136960.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/7d814403-8a1a-482d-afa5-63e68ef4ed09
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/7d814403-8a1a-482d-afa5-63e68ef4ed09",
"name": "7d814403-8a1a-482d-afa5-63e68ef4ed09",
"status": "Failed",
"startTime": "2026-08-24T21:13:13.654576687Z",
"endTime": "2026-08-24T21:32:17.358471293Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable: Waiting for components to be available: router; hosted cluster degraded: UnavailableReplicas: router deployment has 2 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_feature_aggregation.go:184]: failed to create HCP cluster "agg-cluster" with aggregated settings
Unexpected error:
<*fmt.wrapError | 0xc00133a200>:
failed waiting for cluster="agg-cluster" in resourcegroup="feature-aggregation-qp6jkgrm7jtk" to finish creating: GET https://rp.j8136960.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/7d814403-8a1a-482d-afa5-63e68ef4ed09
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/7d814403-8a1a-482d-afa5-63e68ef4ed09",
"name": "7d814403-8a1a-482d-afa5-63e68ef4ed09",
"status": "Failed",
"startTime": "2026-08-24T21:13:13.654576687Z",
"endTime": "2026-08-24T21:32:17.358471293Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable: Waiting for components to be available: router; hosted cluster degraded: UnavailableReplicas: router deployment has 2 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_complex_cilium_kv.go:106]: failed to create HCP cluster "cilium-cluster" with no CNI and private etcd
Unexpected error:
<*fmt.wrapError | 0xc0012a4ea0>:
failed waiting for cluster="cilium-cluster" in resourcegroup="complex-cilium-kv-26mdm7dwkvm7" to finish creating: GET https://rp.j8136960.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/fea58607-6b10-431a-a2a1-cc0c116e10b9
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/fea58607-6b10-431a-a2a1-cc0c116e10b9",
"name": "fea58607-6b10-431a-a2a1-cc0c116e10b9",
"status": "Failed",
"startTime": "2026-08-24T21:12:09.791388432Z",
"endTime": "2026-08-24T21:31:17.457093893Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable: Waiting for components to be available: router; hosted cluster degraded: UnavailableReplicas: router deployment has 2 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_complex_cilium_kv.go:106]: failed to create HCP cluster "cilium-cluster" with no CNI and private etcd
Unexpected error:
<*fmt.wrapError | 0xc0012a4ea0>:
failed waiting for cluster="cilium-cluster" in resourcegroup="complex-cilium-kv-26mdm7dwkvm7" to finish creating: GET https://rp.j8136960.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/fea58607-6b10-431a-a2a1-cc0c116e10b9
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/fea58607-6b10-431a-a2a1-cc0c116e10b9",
"name": "fea58607-6b10-431a-a2a1-cc0c116e10b9",
"status": "Failed",
"startTime": "2026-08-24T21:12:09.791388432Z",
"endTime": "2026-08-24T21:31:17.457093893Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable: Waiting for components to be available: router; hosted cluster degraded: UnavailableReplicas: router deployment has 2 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (2)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Show signal detailsLikely regression — post-good=0; only seen in DEV; only seen in one PRERROR CODE: DeploymentFailed; detail code StorageAccountOperationInProgress; detail message An operation is currently performing on this storage account that requires exclusive ac...full failure pattern: ERROR CODE: DeploymentFailed; detail code StorageAccountOperationInProgress; detail message An operation is currently performing on this storage account that requires exclusive access. Signal: RegressionSignal: Regression — post-good=0; only seen in DEV; only seen in one PR | provision | 1 | 0.97%1 of 103 job runs affected | RegressionSignal: Regression — post-good=0; only seen in DEV; only seen in one PR | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)time=2026-08-24T16:09:49.176Z level=INFO msg="Running step." serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=infra stamp=2 description="Step infra\n Kind: ARM\n Template: templates/mgmt-infra.bicep\n Parameters: configurations/mgmt-infra.tmpl.bicepparam"
time=2026-08-24T16:09:49.840Z level=DEBUG+2 msg="Starting ARM deployment" serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=infra stamp=2
time=2026-08-24T16:09:50.553Z level=DEBUG+3 msg="Deployment started" serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=infra stamp=2 deployment=c783e3e3d697ee4d22dffbab4271479e105f154955a91c3a47e9819df57ac290 portal=https://ms.portal.azure.com/#view/Microsoft_Azure_Resources/DeploymentDetails.MenuView/~/overview/id/%2Fsubscriptions%2FXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX%2FresourceGroups%2Fhcp-underlay-ci01-j1968512-mgmt-2%2Fproviders%2FMicrosoft.Resources%2Fdeployments%2Fc783e3e3d697ee4d22dffbab4271479e105f154955a91c3a47e9819df57ac290
time=2026-08-24T16:11:20.899Z level=ERROR msg="Step errored." serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=infra stamp=2 err="stamp 2: failed to run ARM step: failed to poll deployment: failed to wait for deployment completion: GET https://management.azure.com/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j1968512-mgmt-2/providers/Microsoft.Resources/deployments/c783e3e3d697ee4d22dffbab4271479e105f154955a91c3a47e9819df57ac290/operationStatuses/08584140190952500816\n--------------------------------------------------------------------------------\nRESPONSE 200: 200 OK\nERROR CODE: DeploymentFailed\n--------------------------------------------------------------------------------\n{\n \"status\": \"Failed\",\n \"error\": {\n \"code\": \"DeploymentFailed\",\n \"message\": \"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\",\n \"details\": [\n {\n \"code\": \"Conflict\",\n \"message\": \"{\\r\\n \\\"status\\\": \\\"Failed\\\",\\r\\n \\\"error\\\": {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"DeploymentFailed\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j1968512-mgmt-2/providers/Microsoft.Resources/deployments/hcp-backups-storage\\\",\\r\\n \\\"message\\\": \\\"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j1968512-mgmt-2/providers/Microsoft.Resources/deployments/hcpBackupsStorageAccount\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"DeploymentFailed\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j1968512-mgmt-2/providers/Microsoft.Resources/deployments/hcpBackupsStorageAccount\\\",\\r\\n \\\"message\\\": \\\"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"StorageAccountOperationInProgress\\\",\\r\\n \\\"message\\\": \\\"An operation is currently performing on this storage account that requires exclusive access.\\\"\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n}\"\n }\n ]\n }\n}\n--------------------------------------------------------------------------------\n"Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Show signal detailsLikely regression — post-good=0; only seen in DEV; only seen in one PRERROR CODE: DeploymentFailed; detail code Timeout; detail message Operation is being Canceled due to timeout.; provider Microsoft.EventGridfull failure pattern: ERROR CODE: DeploymentFailed; detail code Timeout; detail message Operation is being Canceled due to timeout.; provider Microsoft.EventGrid Signal: RegressionSignal: Regression — post-good=0; only seen in DEV; only seen in one PR | provision | 1 | 0.97%1 of 103 job runs affected | RegressionSignal: Regression — post-good=0; only seen in DEV; only seen in one PR | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)time=2026-08-24T07:13:06.981Z level=INFO msg="Running step." serviceGroup=Microsoft.Azure.ARO.HCP.Region resourceGroup=regional step=infra description="Step infra\n Kind: ARM\n Template: templates/region.bicep\n Parameters: configurations/region.tmpl.bicepparam"
time=2026-08-24T07:13:09.018Z level=DEBUG+2 msg="Starting ARM deployment" serviceGroup=Microsoft.Azure.ARO.HCP.Region resourceGroup=regional step=infra
time=2026-08-24T07:13:10.536Z level=DEBUG+3 msg="Deployment started" serviceGroup=Microsoft.Azure.ARO.HCP.Region resourceGroup=regional step=infra deployment=9665ebaf931e927bcf7f137293ddd5a9262983f1f2075a0dff43a22749c6dc59 portal=https://ms.portal.azure.com/#view/Microsoft_Azure_Resources/DeploymentDetails.MenuView/~/overview/id/%2Fsubscriptions%2FXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX%2FresourceGroups%2Fhcp-underlay-ci01-j1860736%2Fproviders%2FMicrosoft.Resources%2Fdeployments%2F9665ebaf931e927bcf7f137293ddd5a9262983f1f2075a0dff43a22749c6dc59
time=2026-08-24T07:29:44.382Z level=ERROR msg="Step errored." serviceGroup=Microsoft.Azure.ARO.HCP.Region resourceGroup=regional step=infra err="failed to run ARM step: failed to poll deployment: failed to wait for deployment completion: GET https://management.azure.com/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j1860736/providers/Microsoft.Resources/deployments/9665ebaf931e927bcf7f137293ddd5a9262983f1f2075a0dff43a22749c6dc59/operationStatuses/08584140512958847265\n--------------------------------------------------------------------------------\nRESPONSE 200: 200 OK\nERROR CODE: DeploymentFailed\n--------------------------------------------------------------------------------\n{\n \"status\": \"Failed\",\n \"error\": {\n \"code\": \"DeploymentFailed\",\n \"message\": \"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\",\n \"details\": [\n {\n \"code\": \"Conflict\",\n \"message\": \"{\\r\\n \\\"status\\\": \\\"Failed\\\",\\r\\n \\\"error\\\": {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"DeploymentFailed\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j1860736/providers/Microsoft.Resources/deployments/maestro-infra-deployment\\\",\\r\\n \\\"message\\\": \\\"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j1860736/providers/Microsoft.EventGrid/namespaces/arohcp-ci01-maestro-j1860736\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"Timeout\\\",\\r\\n \\\"message\\\": \\\"Operation is being Canceled due to timeout.\\\"\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n}\"\n }\n ]\n }\n}\n--------------------------------------------------------------------------------\n"Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Show signal detailsLikely regression — post-good=0; only seen in DEV; only seen in one PRERROR CODE: DeploymentStackDeploymentFailed; provider Microsoft.Computefull failure pattern: ERROR CODE: DeploymentStackDeploymentFailed; provider Microsoft.Compute Signal: RegressionSignal: Regression — post-good=0; only seen in DEV; only seen in one PR | provision | 1 | 0.97%1 of 103 job runs affected | RegressionSignal: Regression — post-good=0; only seen in DEV; only seen in one PR | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)time=2026-08-24T08:15:25.803Z level=INFO msg="Running step." serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=nodepools stamp=2 description="Step nodepools\n Kind: ARMStack\n Template: templates/aks-nodepools.bicep\n Parameters: configurations/mgmt-cluster-nodepools.tmpl.bicepparam"
time=2026-08-24T08:18:16.150Z level=ERROR msg="Step errored." serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=nodepools stamp=2 err="stamp 2: failed to run ARM step: failed to wait for deployment stack at resource group scope: GET https://management.azure.com/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.Resources/locations/westus2/deploymentStackOperationStatus/67d8fd4b-85de-4d57-a829-7ce0bf268c5d\n--------------------------------------------------------------------------------\nRESPONSE 200: 200 OK\nERROR CODE: DeploymentStackDeploymentFailed\n--------------------------------------------------------------------------------\n{\n \"id\": \"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.Resources/locations/westus2/deploymentStackOperationStatus/67d8fd4b-85de-4d57-a829-7ce0bf268c5d\",\n \"name\": \"67d8fd4b-85de-4d57-a829-7ce0bf268c5d\",\n \"status\": \"failed\",\n \"error\": {\n \"code\": \"DeploymentStackDeploymentFailed\",\n \"message\": \"One or more resources could not be deployed. Correlation id: '676c3fe4-b0a1-44f5-8602-c3c05012b379'.\",\n \"details\": [\n {\n \"code\": \"DeploymentFailed\",\n \"target\": \"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j9428224-mgmt-2/providers/Microsoft.Resources/deployments/l4mv3scto7z6edigthkjdr62jbdrgoxoif33nvakxlwvg-260824084ipd0\",\n \"message\": \"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\",\n \"details\": [\n {\n \"code\": \"OverconstrainedZonalAllocationRequest\",\n \"message\": \"Allocation failed on resource '/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j9428224-mgmt-2-aks1/providers/Microsoft.Compute/virtualMachineScaleSets/aks-infra1-41975975-vmss'. Please try alternative sizes 'Standard_D8s_v6' (zones 3, 2), 'Standard_E8-4s_v6' (zones 3, 2), 'Standard_E8-4ds_v6' (zones 2, 3) instead for higher likelihood of success. Read more about improving likelihood of operation success at https://aka.ms/AKSallocation. Create or update VMSS /subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j9428224-mgmt-2-aks1/providers/Microsoft.Compute/virtualMachineScaleSets/aks-infra1-41975975-vmss failed. Allocation failed. VM(s) with the following constraints cannot be allocated, because the condition is too restrictive. Please remove some constraints and try again. Constraints applied are:\\n - Availability Zone\\n - Differencing (Ephemeral) Disks\\n - Networking Constraints (such as Accelerated Networking or IPv6)\\n - VM Size\\n\"\n }\n ]\n }\n ]\n }\n}\n--------------------------------------------------------------------------------\n"Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Show signal detailsLikely regression — post-good=0; only seen in DEV; only seen in one PRERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable:...full failure pattern: ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; hosted cluster degraded: UnavailableReplicas: [cluster-version-operator]; provider Microsoft.RedHatOpenShift Signal: RegressionSignal: Regression — post-good=0; only seen in DEV; only seen in one PR | e2e | 1 | 0.97%1 of 103 job runs affected | RegressionSignal: Regression — post-good=0; only seen in DEV; only seen in one PR | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/e2e/hcp_backups.go:85]: failed to create HCP cluster
Unexpected error:
<*fmt.wrapError | 0xc000ef90e0>:
failed to create HCP cluster pause-bkp-cluster: failed waiting for cluster="pause-bkp-cluster" in resourcegroup="pause-bkp-e2e-2zg88lxq9q4w" to finish creating: GET https://rp.j3558656.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/b976f9f3-2891-4433-868a-76a8d23de317
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/b976f9f3-2891-4433-868a-76a8d23de317",
"name": "b976f9f3-2891-4433-868a-76a8d23de317",
"status": "Failed",
"startTime": "2026-08-24T16:54:08.369328068Z",
"endTime": "2026-08-24T17:13:08.208412685Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; hosted cluster degraded: UnavailableReplicas: cluster-version-operator deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/hcp_backups.go:85]: failed to create HCP cluster
Unexpected error:
<*fmt.wrapError | 0xc000ef90e0>:
failed to create HCP cluster pause-bkp-cluster: failed waiting for cluster="pause-bkp-cluster" in resourcegroup="pause-bkp-e2e-2zg88lxq9q4w" to finish creating: GET https://rp.j3558656.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/b976f9f3-2891-4433-868a-76a8d23de317
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/b976f9f3-2891-4433-868a-76a8d23de317",
"name": "b976f9f3-2891-4433-868a-76a8d23de317",
"status": "Failed",
"startTime": "2026-08-24T16:54:08.369328068Z",
"endTime": "2026-08-24T17:13:08.208412685Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; hosted cluster degraded: UnavailableReplicas: cluster-version-operator deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Show signal detailsLikely regression — post-good=0; only seen in DEV; only seen in one PRrun command "e2e-runcommand-1787582391625387873" failed with exit code 1 on VM "<vm>" (output: "Client Version: v1.36.4 Kustomize Version: v5.8.1")full failure pattern: run command "e2e-runcommand-1787582391625387873" failed with exit code 1 on VM "<vm>" (output: "Client Version: v1.36.4 Kustomize Version: v5.8.1") Signal: RegressionSignal: Regression — post-good=0; only seen in DEV; only seen in one PR | e2e | 1 | 0.97%1 of 103 job runs affected | RegressionSignal: Regression — post-good=0; only seen in DEV; only seen in one PR | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_private_kas.go:160]: kubectl version should succeed from VM via private KAS internal LB (output: )
Unexpected error:
<*errors.errorString | 0xc0007fbf60>:
run command "e2e-runcommand-1787582391625387873" failed with exit code 1 on VM "private-kas-test-vm" (output: "Client Version: v1.36.4\nKustomize Version: v5.8.1")
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_private_kas.go:160]: kubectl version should succeed from VM via private KAS internal LB (output: )
Unexpected error:
<*errors.errorString | 0xc0007fbf60>:
run command "e2e-runcommand-1787582391625387873" failed with exit code 1 on VM "private-kas-test-vm" (output: "Client Version: v1.36.4\nKustomize Version: v5.8.1")
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message [hypershiftHostedCluster] hosted cluster has no installed version; hosted cluster degraded: UnavailableReplicas: [router]; provider...full failure pattern: ERROR CODE: InternalServerError; detail message [hypershiftHostedCluster] hosted cluster has no installed version; hosted cluster degraded: UnavailableReplicas: [router]; provider Microsoft.RedHatOpenShift Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | e2e | 20 | 19.42%20 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | Show trend detailsAug 18: 18 · Aug 19: 2 · Aug 20: 3 · Aug 21: 0 · Aug 22: 4 · Aug 23: 0 · Aug 24: 31 | INT | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (3)fail [github.com/Azure/ARO-HCP/test/e2e/complete_cluster_create_multiversion.go:123]: HCP cluster rg-candidate-5-0-h74snzg5tpbh/cluster-candidate-5-0-gp9ztx should provision
Unexpected error:
<*fmt.wrapError | 0xc0014a05c0>:
failed to create HCP cluster cluster-candidate-5-0-gp9ztx: failed waiting for cluster="cluster-candidate-5-0-gp9ztx" in resourcegroup="rg-candidate-5-0-h74snzg5tpbh" to finish creating: GET https://rp.j1843200.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/de4739f4-eebd-4fe2-bf80-3bf6d57ff617
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/de4739f4-eebd-4fe2-bf80-3bf6d57ff617",
"name": "de4739f4-eebd-4fe2-bf80-3bf6d57ff617",
"status": "Failed",
"startTime": "2026-08-25T00:23:47.993778746Z",
"endTime": "2026-08-25T00:42:49.076566651Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster has no installed version; hosted cluster degraded: UnavailableReplicas: router deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/complete_cluster_create_multiversion.go:123]: HCP cluster rg-candidate-5-0-h74snzg5tpbh/cluster-candidate-5-0-gp9ztx should provision
Unexpected error:
<*fmt.wrapError | 0xc0014a05c0>:
failed to create HCP cluster cluster-candidate-5-0-gp9ztx: failed waiting for cluster="cluster-candidate-5-0-gp9ztx" in resourcegroup="rg-candidate-5-0-h74snzg5tpbh" to finish creating: GET https://rp.j1843200.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/de4739f4-eebd-4fe2-bf80-3bf6d57ff617
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/de4739f4-eebd-4fe2-bf80-3bf6d57ff617",
"name": "de4739f4-eebd-4fe2-bf80-3bf6d57ff617",
"status": "Failed",
"startTime": "2026-08-25T00:23:47.993778746Z",
"endTime": "2026-08-25T00:42:49.076566651Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster has no installed version; hosted cluster degraded: UnavailableReplicas: router deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/complete_cluster_create_multiversion.go:123]: HCP cluster rg-candidate-4-21-x45vwlgpgvzk/cluster-candidate-4-21-rztpw4 should provision
Unexpected error:
<*fmt.wrapError | 0xc000f3e6a0>:
failed to create HCP cluster cluster-candidate-4-21-rztpw4: failed waiting for cluster="cluster-candidate-4-21-rztpw4" in resourcegroup="rg-candidate-4-21-x45vwlgpgvzk" to finish creating: GET https://rp.j4709760.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/e986c420-0503-4186-9a01-bad1f92a005c
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/e986c420-0503-4186-9a01-bad1f92a005c",
"name": "e986c420-0503-4186-9a01-bad1f92a005c",
"status": "Failed",
"startTime": "2026-08-24T23:37:07.202301694Z",
"endTime": "2026-08-24T23:56:10.677598487Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster has no installed version; hosted cluster degraded: UnavailableReplicas: router deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/complete_cluster_create_multiversion.go:123]: HCP cluster rg-candidate-4-21-x45vwlgpgvzk/cluster-candidate-4-21-rztpw4 should provision
Unexpected error:
<*fmt.wrapError | 0xc000f3e6a0>:
failed to create HCP cluster cluster-candidate-4-21-rztpw4: failed waiting for cluster="cluster-candidate-4-21-rztpw4" in resourcegroup="rg-candidate-4-21-x45vwlgpgvzk" to finish creating: GET https://rp.j4709760.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/e986c420-0503-4186-9a01-bad1f92a005c
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/e986c420-0503-4186-9a01-bad1f92a005c",
"name": "e986c420-0503-4186-9a01-bad1f92a005c",
"status": "Failed",
"startTime": "2026-08-24T23:37:07.202301694Z",
"endTime": "2026-08-24T23:56:10.677598487Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster has no installed version; hosted cluster degraded: UnavailableReplicas: router deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_private_kas.go:90]: failed to create HCP cluster "private-kas" with private KAS
Unexpected error:
<*fmt.wrapError | 0xc000d48840>:
failed to create HCP cluster private-kas: failed waiting for cluster="private-kas" in resourcegroup="private-kas-mglhphvdtjhd" to finish creating: GET https://rp.j6339200.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/1dbfdf90-d2f5-4efc-b68b-42b49d8220eb
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/1dbfdf90-d2f5-4efc-b68b-42b49d8220eb",
"name": "1dbfdf90-d2f5-4efc-b68b-42b49d8220eb",
"status": "Failed",
"startTime": "2026-08-24T23:29:18.405302537Z",
"endTime": "2026-08-24T23:48:26.458954406Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster has no installed version; hosted cluster degraded: UnavailableReplicas: router deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_private_kas.go:90]: failed to create HCP cluster "private-kas" with private KAS
Unexpected error:
<*fmt.wrapError | 0xc000d48840>:
failed to create HCP cluster private-kas: failed waiting for cluster="private-kas" in resourcegroup="private-kas-mglhphvdtjhd" to finish creating: GET https://rp.j6339200.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/1dbfdf90-d2f5-4efc-b68b-42b49d8220eb
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/1dbfdf90-d2f5-4efc-b68b-42b49d8220eb",
"name": "1dbfdf90-d2f5-4efc-b68b-42b49d8220eb",
"status": "Failed",
"startTime": "2026-08-24T23:29:18.405302537Z",
"endTime": "2026-08-24T23:48:26.458954406Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster has no installed version; hosted cluster degraded: UnavailableReplicas: router deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (18)
Affected runs (20) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [hcp] KubePodNotReady firedfull failure pattern: alert [hcp] KubePodNotReady fired Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 6 days | alert | 14 | 13.59%14 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 6 days | Show trend detailsAug 18: 14 · Aug 19: 9 · Aug 20: 14 · Aug 21: 12 · Aug 22: 2 · Aug 23: 0 · Aug 24: 14 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (3)alert fired 1 time(s) Firing 1: State: Resolved Started: 2026-08-25T00:35:25Z Ended: 2026-08-25T00:56:19Z Severity: Sev3 Labels: alertname="KubePodNotReady", cluster="ci01-j1843200-mgmt-2", component="kubernetes-infrastructure", namespace="ocm-arohcpci01-2sd9551d19b6bblvufm3l3a421htavmi-a4k0z5e8x0e3t0f", pod="router-597d99bc98-tlmg9", severity="warning" Description: Pod ocm-arohcpci01-2sd9551d19b6bblvufm3l3a421htavmi-a4k0z5e8x0e3t0f/router-597d99bc98-tlmg9 has been in a non-ready state for longer than 5 minutes. alert fired 1 time(s) Firing 1: State: Resolved Started: 2026-08-24T23:47:43Z Ended: 2026-08-25T00:09:40Z Severity: Sev3 Labels: alertname="KubePodNotReady", cluster="ci01-j4709760-mgmt-2", component="kubernetes-infrastructure", namespace="ocm-arohcpci01-2sd8f93h71j1qrcah2q03eja41gojmif-i6w1a1c6f6m8f7l", pod="router-7f97cf9c5-phjp6", severity="warning" Description: Pod ocm-arohcpci01-2sd8f93h71j1qrcah2q03eja41gojmif-i6w1a1c6f6m8f7l/router-7f97cf9c5-phjp6 has been in a non-ready state for longer than 5 minutes. alert fired 2 time(s) Firing 1: State: Resolved Started: 2026-08-24T22:04:58Z Ended: 2026-08-24T22:16:56Z Severity: Sev3 Labels: alertname="KubePodNotReady", cluster="ci01-j9409536-mgmt-1", component="kubernetes-infrastructure", namespace="ocm-arohcpci01-2sd6tku01d4e5c0s0k2ocfdtf91cpr8l-d3q6r6q7q0m9f7w", pod="router-57846fcf5b-n5mlw", severity="warning" Description: Pod ocm-arohcpci01-2sd6tku01d4e5c0s0k2ocfdtf91cpr8l-d3q6r6q7q0m9f7w/router-57846fcf5b-n5mlw has been in a non-ready state for longer than 5 minutes. Firing 2: State: Resolved Started: 2026-08-24T22:04:58Z Ended: 2026-08-24T22:16:56Z Severity: Sev3 Labels: alertname="KubePodNotReady", cluster="ci01-j9409536-mgmt-1", component="kubernetes-infrastructure", namespace="ocm-arohcpci01-2sd6td57g5fetk5hnhg543eg0jfi9lrt-j5s2f3c0f5x7m8z", pod="router-7586ffccfd-4mvlg", severity="warning" Description: Pod ocm-arohcpci01-2sd6td57g5fetk5hnhg543eg0jfi9lrt-j5s2f3c0f5x7m8z/router-7586ffccfd-4mvlg has been in a non-ready state for longer than 5 minutes. Contributing tests (1)
Affected runs (14) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
cluster entered terminal Failed state after role assignment deploymentfull failure pattern: cluster entered terminal Failed state after role assignment deployment Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 4 days | e2e | 6 | 5.83%6 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 4 days | Show trend detailsAug 18: 1 · Aug 19: 1 · Aug 20: 0 · Aug 21: 0 · Aug 22: 1 · Aug 23: 0 · Aug 24: 6 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (3)fail [github.com/Azure/ARO-HCP/test/e2e/cluster_delayed_role_assignments.go:227]: Timed out after 1980.001s.
cluster should eventually succeed after role assignments are created
The function passed to Eventually failed at /opt/app-root/src/github.com/Azure/ARO-HCP/test/e2e/cluster_delayed_role_assignments.go:223 with:
cluster entered terminal Failed state after role assignment deployment
Expected
<armredhatopenshifthcp.ProvisioningState>: ...
not to equal
<armredhatopenshifthcp.ProvisioningState>: ...
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_delayed_role_assignments.go:227]: Timed out after 1980.001s.
cluster should eventually succeed after role assignments are created
The function passed to Eventually failed at /opt/app-root/src/github.com/Azure/ARO-HCP/test/e2e/cluster_delayed_role_assignments.go:223 with:
cluster entered terminal Failed state after role assignment deployment
Expected
<armredhatopenshifthcp.ProvisioningState>: ...
not to equal
<armredhatopenshifthcp.ProvisioningState>: ...fail [github.com/Azure/ARO-HCP/test/e2e/cluster_delayed_role_assignments.go:227]: Timed out after 1980.002s.
cluster should eventually succeed after role assignments are created
The function passed to Eventually failed at /opt/app-root/src/github.com/Azure/ARO-HCP/test/e2e/cluster_delayed_role_assignments.go:223 with:
cluster entered terminal Failed state after role assignment deployment
Expected
<armredhatopenshifthcp.ProvisioningState>: ...
not to equal
<armredhatopenshifthcp.ProvisioningState>: ...
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_delayed_role_assignments.go:227]: Timed out after 1980.002s.
cluster should eventually succeed after role assignments are created
The function passed to Eventually failed at /opt/app-root/src/github.com/Azure/ARO-HCP/test/e2e/cluster_delayed_role_assignments.go:223 with:
cluster entered terminal Failed state after role assignment deployment
Expected
<armredhatopenshifthcp.ProvisioningState>: ...
not to equal
<armredhatopenshifthcp.ProvisioningState>: ...fail [github.com/Azure/ARO-HCP/test/e2e/cluster_delayed_role_assignments.go:227]: Timed out after 1980.000s.
cluster should eventually succeed after role assignments are created
The function passed to Eventually failed at /opt/app-root/src/github.com/Azure/ARO-HCP/test/e2e/cluster_delayed_role_assignments.go:223 with:
cluster entered terminal Failed state after role assignment deployment
Expected
<armredhatopenshifthcp.ProvisioningState>: ...
not to equal
<armredhatopenshifthcp.ProvisioningState>: ...
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_delayed_role_assignments.go:227]: Timed out after 1980.000s.
cluster should eventually succeed after role assignments are created
The function passed to Eventually failed at /opt/app-root/src/github.com/Azure/ARO-HCP/test/e2e/cluster_delayed_role_assignments.go:223 with:
cluster entered terminal Failed state after role assignment deployment
Expected
<armredhatopenshifthcp.ProvisioningState>: ...
not to equal
<armredhatopenshifthcp.ProvisioningState>: ...Contributing tests (1)
Affected runs (6) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; provider Microsoft.RedHatOpenShiftfull failure pattern: ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; provider Microsoft.RedHatOpenShift Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | e2e | 4 | 3.88%4 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | Show trend detailsAug 18: 3 · Aug 19: 2 · Aug 20: 0 · Aug 21: 27 · Aug 22: 5 · Aug 23: 0 · Aug 24: 4 | STG | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (3)fail [github.com/Azure/ARO-HCP/test/e2e/cluster_autoscaling.go:109]: failed to create HCP cluster "autoscaling-hcp-cluster" with custom autoscaling
Unexpected error:
<*fmt.wrapError | 0xc000da71c0>:
failed to create HCP cluster autoscaling-hcp-cluster: failed waiting for cluster="autoscaling-hcp-cluster" in resourcegroup="autoscaling-cluster-wqnnn79bk4qc" to finish creating: GET https://rp.j6887040.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/e635655f-c9ac-44e1-8232-3550a4c20673
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/e635655f-c9ac-44e1-8232-3550a4c20673",
"name": "e635655f-c9ac-44e1-8232-3550a4c20673",
"status": "Failed",
"startTime": "2026-08-24T21:28:57.457434277Z",
"endTime": "2026-08-24T21:48:06.423004565Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_autoscaling.go:109]: failed to create HCP cluster "autoscaling-hcp-cluster" with custom autoscaling
Unexpected error:
<*fmt.wrapError | 0xc000da71c0>:
failed to create HCP cluster autoscaling-hcp-cluster: failed waiting for cluster="autoscaling-hcp-cluster" in resourcegroup="autoscaling-cluster-wqnnn79bk4qc" to finish creating: GET https://rp.j6887040.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/e635655f-c9ac-44e1-8232-3550a4c20673
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/e635655f-c9ac-44e1-8232-3550a4c20673",
"name": "e635655f-c9ac-44e1-8232-3550a4c20673",
"status": "Failed",
"startTime": "2026-08-24T21:28:57.457434277Z",
"endTime": "2026-08-24T21:48:06.423004565Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e"
}
}
--------------------------------------------------------------------------------
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/cluster_autoscaling.go:109]: failed to create HCP cluster "autoscaling-hcp-cluster" with custom autoscaling
Unexpected error:
<*fmt.wrapError | 0xc000fa81a0>:
failed to create HCP cluster autoscaling-hcp-cluster: failed waiting for cluster="autoscaling-hcp-cluster" in resourcegroup="autoscaling-cluster-jhvt8ptm7fw8" to finish creating: GET https://rp.j8539776.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/12485de2-db77-439f-89df-5bacb5d218b2
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/12485de2-db77-439f-89df-5bacb5d218b2",
"name": "12485de2-db77-439f-89df-5bacb5d218b2",
"status": "Failed",
"startTime": "2026-08-24T15:41:38.67903036Z",
"endTime": "2026-08-24T16:00:39.109568382Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_autoscaling.go:109]: failed to create HCP cluster "autoscaling-hcp-cluster" with custom autoscaling
Unexpected error:
<*fmt.wrapError | 0xc000fa81a0>:
failed to create HCP cluster autoscaling-hcp-cluster: failed waiting for cluster="autoscaling-hcp-cluster" in resourcegroup="autoscaling-cluster-jhvt8ptm7fw8" to finish creating: GET https://rp.j8539776.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/12485de2-db77-439f-89df-5bacb5d218b2
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/12485de2-db77-439f-89df-5bacb5d218b2",
"name": "12485de2-db77-439f-89df-5bacb5d218b2",
"status": "Failed",
"startTime": "2026-08-24T15:41:38.67903036Z",
"endTime": "2026-08-24T16:00:39.109568382Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e"
}
}
--------------------------------------------------------------------------------
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/cluster_update.go:161]: failed to create HCP cluster for patch-tags test
Unexpected error:
<*fmt.wrapError | 0xc000e86920>:
failed to create HCP cluster patch-tags-cluster: failed waiting for cluster="patch-tags-cluster" in resourcegroup="patch-tags-7vpnvrbqnldf" to finish creating: GET https://rp.j1362816.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/73cc6203-30af-4c53-9024-398221683c83
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/73cc6203-30af-4c53-9024-398221683c83",
"name": "73cc6203-30af-4c53-9024-398221683c83",
"status": "Failed",
"startTime": "2026-08-24T15:17:58.555796542Z",
"endTime": "2026-08-24T15:37:08.01693151Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_update.go:161]: failed to create HCP cluster for patch-tags test
Unexpected error:
<*fmt.wrapError | 0xc000e86920>:
failed to create HCP cluster patch-tags-cluster: failed waiting for cluster="patch-tags-cluster" in resourcegroup="patch-tags-7vpnvrbqnldf" to finish creating: GET https://rp.j1362816.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/73cc6203-30af-4c53-9024-398221683c83
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/73cc6203-30af-4c53-9024-398221683c83",
"name": "73cc6203-30af-4c53-9024-398221683c83",
"status": "Failed",
"startTime": "2026-08-24T15:17:58.555796542Z",
"endTime": "2026-08-24T15:37:08.01693151Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (3)
Affected runs (4)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster has no installed version; provider Microsoft.R...full failure pattern: ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster has no installed version; provider Microsoft.RedHatOpenShift Signal: FlakeSignal: Flake — present in 2 prior week(s); spread across 6 days | e2e | 4 | 3.88%4 of 103 job runs affected | FlakeSignal: Flake — present in 2 prior week(s); spread across 6 days | Show trend detailsAug 18: 3 · Aug 19: 2 · Aug 20: 1 · Aug 21: 1 · Aug 22: 1 · Aug 23: 0 · Aug 24: 4 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (3)fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_nodepool_osdisk.go:80]: failed to create HCP cluster "hcp-cluster-np-128"
Unexpected error:
<*fmt.wrapError | 0xc00123e240>:
failed to create HCP cluster hcp-cluster-np-128: failed waiting for cluster="hcp-cluster-np-128" in resourcegroup="clusternp128-2xcnpkr48kxk" to finish creating: GET https://rp.j2463616.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/3185d8c9-e246-4082-b170-898da28192d1
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/3185d8c9-e246-4082-b170-898da28192d1",
"name": "3185d8c9-e246-4082-b170-898da28192d1",
"status": "Failed",
"startTime": "2026-08-24T17:28:32.888311428Z",
"endTime": "2026-08-24T17:47:34.534131157Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster has no installed version"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_nodepool_osdisk.go:80]: failed to create HCP cluster "hcp-cluster-np-128"
Unexpected error:
<*fmt.wrapError | 0xc00123e240>:
failed to create HCP cluster hcp-cluster-np-128: failed waiting for cluster="hcp-cluster-np-128" in resourcegroup="clusternp128-2xcnpkr48kxk" to finish creating: GET https://rp.j2463616.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/3185d8c9-e246-4082-b170-898da28192d1
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/3185d8c9-e246-4082-b170-898da28192d1",
"name": "3185d8c9-e246-4082-b170-898da28192d1",
"status": "Failed",
"startTime": "2026-08-24T17:28:32.888311428Z",
"endTime": "2026-08-24T17:47:34.534131157Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster has no installed version"
}
}
--------------------------------------------------------------------------------
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/nodepool_version_upgrade.go:726]: failed to create HCP cluster
Unexpected error:
<*fmt.wrapError | 0xc000dd0320>:
failed to create HCP cluster np-dg-minor-4tds4q: failed waiting for cluster="np-dg-minor-4tds4q" in resourcegroup="rg-np-dg-minor-4tds4q-6qtw9l2fwcdn" to finish creating: GET https://rp.j8539776.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/162a10b5-886b-40a8-a95b-30fc38ce9bc5
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/162a10b5-886b-40a8-a95b-30fc38ce9bc5",
"name": "162a10b5-886b-40a8-a95b-30fc38ce9bc5",
"status": "Failed",
"startTime": "2026-08-24T15:41:41.593852701Z",
"endTime": "2026-08-24T16:00:49.106385261Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster has no installed version"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/nodepool_version_upgrade.go:726]: failed to create HCP cluster
Unexpected error:
<*fmt.wrapError | 0xc000dd0320>:
failed to create HCP cluster np-dg-minor-4tds4q: failed waiting for cluster="np-dg-minor-4tds4q" in resourcegroup="rg-np-dg-minor-4tds4q-6qtw9l2fwcdn" to finish creating: GET https://rp.j8539776.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/162a10b5-886b-40a8-a95b-30fc38ce9bc5
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/162a10b5-886b-40a8-a95b-30fc38ce9bc5",
"name": "162a10b5-886b-40a8-a95b-30fc38ce9bc5",
"status": "Failed",
"startTime": "2026-08-24T15:41:41.593852701Z",
"endTime": "2026-08-24T16:00:49.106385261Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster has no installed version"
}
}
--------------------------------------------------------------------------------
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/image_registry_cluster_create.go:89]: failed to create HCP cluster with image registry disabled
Unexpected error:
<*fmt.wrapError | 0xc000e96560>:
failed to create HCP cluster disabled-image-registry-hcp-cluster: failed waiting for cluster="disabled-image-registry-hcp-cluster" in resourcegroup="disabled-image-registry-fn2m5nmfttf7" to finish creating: GET https://rp.j1362816.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/eb7b7e77-c07f-48f4-a1bf-2f48bcb6434a
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/eb7b7e77-c07f-48f4-a1bf-2f48bcb6434a",
"name": "eb7b7e77-c07f-48f4-a1bf-2f48bcb6434a",
"status": "Failed",
"startTime": "2026-08-24T15:17:08.176674236Z",
"endTime": "2026-08-24T15:36:17.96458443Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster has no installed version"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/image_registry_cluster_create.go:89]: failed to create HCP cluster with image registry disabled
Unexpected error:
<*fmt.wrapError | 0xc000e96560>:
failed to create HCP cluster disabled-image-registry-hcp-cluster: failed waiting for cluster="disabled-image-registry-hcp-cluster" in resourcegroup="disabled-image-registry-fn2m5nmfttf7" to finish creating: GET https://rp.j1362816.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/eb7b7e77-c07f-48f4-a1bf-2f48bcb6434a
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/eb7b7e77-c07f-48f4-a1bf-2f48bcb6434a",
"name": "eb7b7e77-c07f-48f4-a1bf-2f48bcb6434a",
"status": "Failed",
"startTime": "2026-08-24T15:17:08.176674236Z",
"endTime": "2026-08-24T15:36:17.96458443Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster has no installed version"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (3)
Affected runs (4)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [hcp] KubeStatefulSetReplicasMismatch firedfull failure pattern: alert [hcp] KubeStatefulSetReplicasMismatch fired Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 6 days | alert | 2 | 1.94%2 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 6 days | Show trend detailsAug 18: 10 · Aug 19: 5 · Aug 20: 6 · Aug 21: 5 · Aug 22: 3 · Aug 23: 0 · Aug 24: 2 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (2)alert fired 2 time(s) Firing 1: State: Resolved Started: 2026-08-24T18:01:25Z Ended: 2026-08-24T18:34:22Z Severity: Sev3 Labels: alertname="KubeStatefulSetReplicasMismatch", cluster="ci01-j0375808-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", endpoint="http", environment="ci01", instance="10.128.64.140:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0375808/providers/microsoft.monitor/accounts/hcps-j0375808", namespace="ocm-arohcpci01-2sd37mkptl2qprba4ejkvlbldgl77u9k-r3i6h6d2n5v5b4i", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-hf5qz", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning", statefulset="etcd" Description: StatefulSet ocm-arohcpci01-2sd37mkptl2qprba4ejkvlbldgl77u9k-r3i6h6d2n5v5b4i/etcd has not matched the expected number of replicas for longer than 15 minutes. Firing 2: State: Resolved Started: 2026-08-24T18:01:25Z Ended: 2026-08-24T18:34:22Z Severity: Sev3 Labels: alertname="KubeStatefulSetReplicasMismatch", cluster="ci01-j0375808-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", endpoint="http", environment="ci01", instance="10.128.64.140:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0375808/providers/microsoft.monitor/accounts/hcps-j0375808", namespace="ocm-arohcpci01-2sd37mkptl2qprba4ejkvlbldgl77u9k-r3i6h6d2n5v5b4i", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-hf5qz", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning", statefulset="etcd" Description: StatefulSet ocm-arohcpci01-2sd37mkptl2qprba4ejkvlbldgl77u9k-r3i6h6d2n5v5b4i/etcd has not matched the expected number of replicas for longer than 15 minutes. alert fired 2 time(s) Firing 1: State: Resolved Started: 2026-08-24T17:55:40Z Ended: 2026-08-24T18:14:40Z Severity: Sev3 Labels: alertname="KubeStatefulSetReplicasMismatch", cluster="ci01-j0629888-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", endpoint="http", environment="ci01", instance="10.128.64.209:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0629888/providers/microsoft.monitor/accounts/hcps-j0629888", namespace="ocm-arohcpci01-2sd354mcvucng8mm6igmcq4k43caf9jb-s0d4l4d8t6j0p0t", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-tktc6", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning", statefulset="etcd" Description: StatefulSet ocm-arohcpci01-2sd354mcvucng8mm6igmcq4k43caf9jb-s0d4l4d8t6j0p0t/etcd has not matched the expected number of replicas for longer than 15 minutes. Firing 2: State: Resolved Started: 2026-08-24T17:55:40Z Ended: 2026-08-24T18:15:41Z Severity: Sev3 Labels: alertname="KubeStatefulSetReplicasMismatch", cluster="ci01-j0629888-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", endpoint="http", environment="ci01", instance="10.128.64.209:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0629888/providers/microsoft.monitor/accounts/hcps-j0629888", namespace="ocm-arohcpci01-2sd354mcvucng8mm6igmcq4k43caf9jb-s0d4l4d8t6j0p0t", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-tktc6", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning", statefulset="etcd" Description: StatefulSet ocm-arohcpci01-2sd354mcvucng8mm6igmcq4k43caf9jb-s0d4l4d8t6j0p0t/etcd has not matched the expected number of replicas for longer than 15 minutes. Contributing tests (1)
Affected runs (2)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message [clusterServiceNodePoolStatus] <no_message>; provider Microsoft.RedHatOpenShiftfull failure pattern: ERROR CODE: InternalServerError; detail message [clusterServiceNodePoolStatus] <no_message>; provider Microsoft.RedHatOpenShift Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 6 days | e2e | 2 | 1.94%2 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 6 days | Show trend detailsAug 18: 2 · Aug 19: 1 · Aug 20: 7 · Aug 21: 2 · Aug 22: 1 · Aug 23: 0 · Aug 24: 2 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (2)fail [github.com/Azure/ARO-HCP/test/e2e/nodepool_version_upgrade.go:166]: failed to create node pool npupgrade-4-20 with version 4.20.34
Unexpected error:
<*fmt.wrapError | 0xc00031f900>:
failed to create NodePool npupgrade-4-20: failed waiting for nodepool="npupgrade-4-20" for cluster "np-version-upgrade-cluster-tczrqg" in resourcegroup="rg-np-version-upgrade-tczrqg-fhlnbg67qjlz" to finish creating: GET https://rp.j7254272.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/c2d7b313-a7f0-40de-93b0-918f61dd2377
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/c2d7b313-a7f0-40de-93b0-918f61dd2377",
"name": "c2d7b313-a7f0-40de-93b0-918f61dd2377",
"status": "Failed",
"startTime": "2026-08-24T20:58:35.24134299Z",
"endTime": "2026-08-24T21:17:39.53487523Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceNodePoolStatus] \u003cno_message\u003e"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/nodepool_version_upgrade.go:166]: failed to create node pool npupgrade-4-20 with version 4.20.34
Unexpected error:
<*fmt.wrapError | 0xc00031f900>:
failed to create NodePool npupgrade-4-20: failed waiting for nodepool="npupgrade-4-20" for cluster "np-version-upgrade-cluster-tczrqg" in resourcegroup="rg-np-version-upgrade-tczrqg-fhlnbg67qjlz" to finish creating: GET https://rp.j7254272.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/c2d7b313-a7f0-40de-93b0-918f61dd2377
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/c2d7b313-a7f0-40de-93b0-918f61dd2377",
"name": "c2d7b313-a7f0-40de-93b0-918f61dd2377",
"status": "Failed",
"startTime": "2026-08-24T20:58:35.24134299Z",
"endTime": "2026-08-24T21:17:39.53487523Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceNodePoolStatus] \u003cno_message\u003e"
}
}
--------------------------------------------------------------------------------
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/nodepool_version_upgrade.go:599]: failed to create nodepool npdg-4-21
Unexpected error:
<*fmt.wrapError | 0xc000e96960>:
failed to create NodePool npdg-4-21: failed waiting for nodepool="npdg-4-21" for cluster "np-downgrade-q2tthc" in resourcegroup="rg-np-downgrade-q2tthc-z7rbl7r2jwws" to finish creating: GET https://rp.j1362816.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/47831462-3285-4a15-8722-8601b5708cf5
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/47831462-3285-4a15-8722-8601b5708cf5",
"name": "47831462-3285-4a15-8722-8601b5708cf5",
"status": "Failed",
"startTime": "2026-08-24T15:29:26.286804415Z",
"endTime": "2026-08-24T15:48:34.685207308Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceNodePoolStatus] \u003cno_message\u003e"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/nodepool_version_upgrade.go:599]: failed to create nodepool npdg-4-21
Unexpected error:
<*fmt.wrapError | 0xc000e96960>:
failed to create NodePool npdg-4-21: failed waiting for nodepool="npdg-4-21" for cluster "np-downgrade-q2tthc" in resourcegroup="rg-np-downgrade-q2tthc-z7rbl7r2jwws" to finish creating: GET https://rp.j1362816.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/47831462-3285-4a15-8722-8601b5708cf5
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/47831462-3285-4a15-8722-8601b5708cf5",
"name": "47831462-3285-4a15-8722-8601b5708cf5",
"status": "Failed",
"startTime": "2026-08-24T15:29:26.286804415Z",
"endTime": "2026-08-24T15:48:34.685207308Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceNodePoolStatus] \u003cno_message\u003e"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (2)
Affected runs (2)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: DeploymentNotFoundfull failure pattern: ERROR CODE: DeploymentNotFound Signal: FlakeSignal: Flake — present in 3 prior week(s); spread across 5 days | provision | 2 | 1.94%2 of 103 job runs affected | FlakeSignal: Flake — present in 3 prior week(s); spread across 5 days | Show trend detailsAug 18: 3 · Aug 19: 1 · Aug 20: 1 · Aug 21: 4 · Aug 22: 0 · Aug 23: 0 · Aug 24: 2 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (2)time=2026-08-24T18:57:37.594Z level=INFO msg="Running step." serviceGroup=Microsoft.Azure.ARO.HCP.Velero resourceGroup=management step=output stamp=1 description="Step output\n Kind: ARM\n Template: ./../dev-infrastructure/templates/output-mgmt.bicep\n Parameters: ./../dev-infrastructure/configurations/output-mgmt.tmpl.bicepparam"
time=2026-08-24T18:57:37.979Z level=DEBUG+2 msg="Starting ARM deployment" serviceGroup=Microsoft.Azure.ARO.HCP.Velero resourceGroup=management step=output stamp=1
time=2026-08-24T18:57:38.350Z level=DEBUG+3 msg="Deployment started" serviceGroup=Microsoft.Azure.ARO.HCP.Velero resourceGroup=management step=output stamp=1 deployment=cfd650f4557893b92396842785417ac4f900952cc40f35ef046b64198c66a837 portal=https://ms.portal.azure.com/#view/Microsoft_Azure_Resources/DeploymentDetails.MenuView/~/overview/id/%2Fsubscriptions%2FXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX%2FresourceGroups%2Fhcp-underlay-ci01-j8797056-mgmt-1%2Fproviders%2FMicrosoft.Resources%2Fdeployments%2Fcfd650f4557893b92396842785417ac4f900952cc40f35ef046b64198c66a837
time=2026-08-24T18:57:38.400Z level=ERROR msg="Step errored." serviceGroup=Microsoft.Azure.ARO.HCP.Velero resourceGroup=management step=output stamp=1 err="stamp 1: failed to run ARM step: failed to poll deployment: failed to wait for deployment completion: GET https://management.azure.com/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j8797056-mgmt-1/providers/Microsoft.Resources/deployments/cfd650f4557893b92396842785417ac4f900952cc40f35ef046b64198c66a837/operationStatuses/08584140090272578417\n--------------------------------------------------------------------------------\nRESPONSE 404: 404 Not Found\nERROR CODE: DeploymentNotFound\n--------------------------------------------------------------------------------\n{\n \"error\": {\n \"code\": \"DeploymentNotFound\",\n \"message\": \"Deployment 'cfd650f4557893b92396842785417ac4f900952cc40f35ef046b64198c66a837' could not be found.\"\n }\n}\n--------------------------------------------------------------------------------\n"time=2026-08-24T17:20:28.034Z level=INFO msg="Running step." serviceGroup=Microsoft.Azure.ARO.HCP.Fleet.Registration resourceGroup=management step=output stamp=1 description="Step output\n Kind: ARM\n Template: ../../dev-infrastructure/templates/output-mgmt.bicep\n Parameters: ../../dev-infrastructure/configurations/output-mgmt.tmpl.bicepparam"
time=2026-08-24T17:20:28.426Z level=DEBUG+2 msg="Starting ARM deployment" serviceGroup=Microsoft.Azure.ARO.HCP.Fleet.Registration resourceGroup=management step=output stamp=1
time=2026-08-24T17:20:28.817Z level=DEBUG+3 msg="Deployment started" serviceGroup=Microsoft.Azure.ARO.HCP.Fleet.Registration resourceGroup=management step=output stamp=1 deployment=3472b226dd1bf3f5c2c2fbd8f62937b0fe2f97cb4ca8a63168999f9ba0945d23 portal=https://ms.portal.azure.com/#view/Microsoft_Azure_Resources/DeploymentDetails.MenuView/~/overview/id/%2Fsubscriptions%2FXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX%2FresourceGroups%2Fhcp-underlay-ci01-j2076032-mgmt-1%2Fproviders%2FMicrosoft.Resources%2Fdeployments%2F3472b226dd1bf3f5c2c2fbd8f62937b0fe2f97cb4ca8a63168999f9ba0945d23
time=2026-08-24T17:20:28.866Z level=ERROR msg="Step errored." serviceGroup=Microsoft.Azure.ARO.HCP.Fleet.Registration resourceGroup=management step=output stamp=1 err="stamp 1: failed to run ARM step: failed to poll deployment: failed to wait for deployment completion: GET https://management.azure.com/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j2076032-mgmt-1/providers/Microsoft.Resources/deployments/3472b226dd1bf3f5c2c2fbd8f62937b0fe2f97cb4ca8a63168999f9ba0945d23/operationStatuses/08584140148567905734\n--------------------------------------------------------------------------------\nRESPONSE 404: 404 Not Found\nERROR CODE: DeploymentNotFound\n--------------------------------------------------------------------------------\n{\n \"error\": {\n \"code\": \"DeploymentNotFound\",\n \"message\": \"Deployment '3472b226dd1bf3f5c2c2fbd8f62937b0fe2f97cb4ca8a63168999f9ba0945d23' could not be found.\"\n }\n}\n--------------------------------------------------------------------------------\n"Contributing tests (2)
Affected runs (2)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message cluster deletion did not complete before the deadline; [clusterServiceDeletion] ClusterService cluster still exists; [clusterServic...full failure pattern: ERROR CODE: InternalServerError; detail message cluster deletion did not complete before the deadline; [clusterServiceDeletion] ClusterService cluster still exists; [clusterServiceStatus] state is "uninstalling"; [descendantResources] remaining resources: serviceProviderClusters; provider Microsoft.RedHatOpenShift Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | e2e | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | Show trend detailsAug 18: 43 · Aug 19: 30 · Aug 20: 18 · Aug 21: 17 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/util/framework/per_test_framework.go:293]: Unexpected error:
<*errors.joinError | 0xc000e8af48>:
failed to cleanup resource group: at least one hcp cluster failed to delete: failed waiting for hcpCluster="sre-hcp-cluster" in resourcegroup="admin-api-breakglass-mb2sm98hvbm8" to finish deleting: GET https://rp.j9012096.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/76928be7-64f8-4a4c-85f6-cd6b790148b9
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/76928be7-64f8-4a4c-85f6-cd6b790148b9",
"name": "76928be7-64f8-4a4c-85f6-cd6b790148b9",
"status": "Failed",
"startTime": "2026-08-24T15:14:43.826676557Z",
"endTime": "2026-08-24T15:38:50.051171369Z",
"error": {
"code": "InternalServerError",
"message": "cluster deletion did not complete before the deadline; [clusterServiceDeletion] ClusterService cluster /api/aro_hcp/v1alpha1/clusters/2sd0tm25j05l3q2f8a1e83cjmn4lda8g still exists (deletion dispatched at 2026-08-24T15:14:45Z); [clusterServiceStatus] ClusterService state is \"uninstalling\"; [descendantResources] remaining resources: 1 microsoft.redhatopenshift/hcpopenshiftclusters/serviceProviderClusters"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/util/framework/per_test_framework.go:293]: Unexpected error:
<*errors.joinError | 0xc000e8af48>:
failed to cleanup resource group: at least one hcp cluster failed to delete: failed waiting for hcpCluster="sre-hcp-cluster" in resourcegroup="admin-api-breakglass-mb2sm98hvbm8" to finish deleting: GET https://rp.j9012096.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/76928be7-64f8-4a4c-85f6-cd6b790148b9
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/76928be7-64f8-4a4c-85f6-cd6b790148b9",
"name": "76928be7-64f8-4a4c-85f6-cd6b790148b9",
"status": "Failed",
"startTime": "2026-08-24T15:14:43.826676557Z",
"endTime": "2026-08-24T15:38:50.051171369Z",
"error": {
"code": "InternalServerError",
"message": "cluster deletion did not complete before the deadline; [clusterServiceDeletion] ClusterService cluster /api/aro_hcp/v1alpha1/clusters/2sd0tm25j05l3q2f8a1e83cjmn4lda8g still exists (deletion dispatched at 2026-08-24T15:14:45Z); [clusterServiceStatus] ClusterService state is \"uninstalling\"; [descendantResources] remaining resources: 1 microsoft.redhatopenshift/hcpopenshiftclusters/serviceProviderClusters"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: KubeconfigWaitingForCreate: ...full failure pattern: ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: KubeconfigWaitingForCreate: Waiting for hosted control plane kubeconfig to be created; hosted cluster degraded: UnavailableReplicas: [capi-provider]; provider Microsoft.RedHatOpenShift Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 6 days | e2e | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 6 days | Show trend detailsAug 18: 26 · Aug 19: 14 · Aug 20: 33 · Aug 21: 40 · Aug 22: 14 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/e2e/complete_cluster_create_multiversion.go:123]: HCP cluster rg-candidate-4-21-nzmkwtbzb8xk/cluster-candidate-4-21-k47dlw should provision
Unexpected error:
<*fmt.wrapError | 0xc00011d300>:
failed to create HCP cluster cluster-candidate-4-21-k47dlw: failed waiting for cluster="cluster-candidate-4-21-k47dlw" in resourcegroup="rg-candidate-4-21-nzmkwtbzb8xk" to finish creating: GET https://rp.j0629888.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/e224492a-2630-4a37-b5e0-7ef00f6cd9dd
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/e224492a-2630-4a37-b5e0-7ef00f6cd9dd",
"name": "e224492a-2630-4a37-b5e0-7ef00f6cd9dd",
"status": "Failed",
"startTime": "2026-08-24T17:34:09.560283017Z",
"endTime": "2026-08-24T17:53:13.775593008Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: KubeconfigWaitingForCreate: Waiting for hosted control plane kubeconfig to be created; hosted cluster degraded: UnavailableReplicas: capi-provider deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/complete_cluster_create_multiversion.go:123]: HCP cluster rg-candidate-4-21-nzmkwtbzb8xk/cluster-candidate-4-21-k47dlw should provision
Unexpected error:
<*fmt.wrapError | 0xc00011d300>:
failed to create HCP cluster cluster-candidate-4-21-k47dlw: failed waiting for cluster="cluster-candidate-4-21-k47dlw" in resourcegroup="rg-candidate-4-21-nzmkwtbzb8xk" to finish creating: GET https://rp.j0629888.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/e224492a-2630-4a37-b5e0-7ef00f6cd9dd
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/e224492a-2630-4a37-b5e0-7ef00f6cd9dd",
"name": "e224492a-2630-4a37-b5e0-7ef00f6cd9dd",
"status": "Failed",
"startTime": "2026-08-24T17:34:09.560283017Z",
"endTime": "2026-08-24T17:53:13.775593008Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: KubeconfigWaitingForCreate: Waiting for hosted control plane kubeconfig to be created; hosted cluster degraded: UnavailableReplicas: capi-provider deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [svc] KubeDaemonSetRolloutStuck firedfull failure pattern: alert [svc] KubeDaemonSetRolloutStuck fired Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | alert | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | Show trend detailsAug 18: 15 · Aug 19: 12 · Aug 20: 10 · Aug 21: 11 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)alert fired 24 time(s) Firing 1: State: Resolved Started: 2026-08-24T16:11:08Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="aks-secrets-store-csi-driver", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/aks-secrets-store-csi-driver has not finished or progressed for at least 15 minutes. Firing 2: State: Resolved Started: 2026-08-24T16:12:09Z Ended: 2026-08-24T16:24:06Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="aks-secrets-store-csi-driver", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/aks-secrets-store-csi-driver has not finished or progressed for at least 15 minutes. Firing 3: State: Resolved Started: 2026-08-24T16:12:09Z Ended: 2026-08-24T16:24:06Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="kube-proxy", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/kube-proxy has not finished or progressed for at least 15 minutes. Firing 4: State: Resolved Started: 2026-08-24T16:12:09Z Ended: 2026-08-24T16:24:06Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="cloud-node-manager", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/cloud-node-manager has not finished or progressed for at least 15 minutes. Firing 5: State: Resolved Started: 2026-08-24T16:12:09Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="csi-azuredisk-node", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/csi-azuredisk-node has not finished or progressed for at least 15 minutes. Firing 6: State: Resolved Started: 2026-08-24T16:12:09Z Ended: 2026-08-24T16:24:06Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="azure-cns", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/azure-cns has not finished or progressed for at least 15 minutes. Firing 7: State: Resolved Started: 2026-08-24T16:12:09Z Ended: 2026-08-24T16:25:05Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="retina-agent", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/retina-agent has not finished or progressed for at least 15 minutes. Firing 8: State: Resolved Started: 2026-08-24T16:12:09Z Ended: 2026-08-24T16:24:06Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="ama-metrics-node", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/ama-metrics-node has not finished or progressed for at least 15 minutes. Firing 9: State: Resolved Started: 2026-08-24T16:12:09Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="csi-azurefile-node", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/csi-azurefile-node has not finished or progressed for at least 15 minutes. Firing 10: State: Resolved Started: 2026-08-24T16:12:09Z Ended: 2026-08-24T16:24:06Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="aks-secrets-store-provider-azure", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/aks-secrets-store-provider-azure has not finished or progressed for at least 15 minutes. Firing 11: State: Resolved Started: 2026-08-24T16:12:09Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="azure-npm", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/azure-npm has not finished or progressed for at least 15 minutes. Firing 12: State: Resolved Started: 2026-08-24T16:13:07Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="aks-secrets-store-provider-azure", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/aks-secrets-store-provider-azure has not finished or progressed for at least 15 minutes. Firing 13: State: Resolved Started: 2026-08-24T16:13:07Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="csi-azurefile-node", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/csi-azurefile-node has not finished or progressed for at least 15 minutes. Firing 14: State: Resolved Started: 2026-08-24T16:13:10Z Ended: 2026-08-24T16:24:06Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="azure-cns", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/azure-cns has not finished or progressed for at least 15 minutes. Firing 15: State: Resolved Started: 2026-08-24T16:13:10Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="csi-azuredisk-node", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/csi-azuredisk-node has not finished or progressed for at least 15 minutes. Firing 16: State: Resolved Started: 2026-08-24T16:13:10Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="kube-proxy", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/kube-proxy has not finished or progressed for at least 15 minutes. Firing 17: State: Resolved Started: 2026-08-24T16:13:10Z Ended: 2026-08-24T16:26:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="retina-agent", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/retina-agent has not finished or progressed for at least 15 minutes. Firing 18: State: Resolved Started: 2026-08-24T16:13:10Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="azure-npm", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/azure-npm has not finished or progressed for at least 15 minutes. Firing 19: State: Resolved Started: 2026-08-24T16:13:10Z Ended: 2026-08-24T16:24:06Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="ama-metrics-node", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/ama-metrics-node has not finished or progressed for at least 15 minutes. Firing 20: State: Resolved Started: 2026-08-24T16:13:10Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="cloud-node-manager", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="kube-system", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet kube-system/cloud-node-manager has not finished or progressed for at least 15 minutes. Firing 21: State: Resolved Started: 2026-08-24T16:17:09Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="arobit-forwarder", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="arobit", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet arobit/arobit-forwarder has not finished or progressed for at least 15 minutes. Firing 22: State: Resolved Started: 2026-08-24T16:17:09Z Ended: 2026-08-24T16:23:05Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="node-agent", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="velero", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet velero/node-agent has not finished or progressed for at least 15 minutes. Firing 23: State: Resolved Started: 2026-08-24T16:17:09Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="node-agent", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="velero", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet velero/node-agent has not finished or progressed for at least 15 minutes. Firing 24: State: Resolved Started: 2026-08-24T16:17:09Z Ended: 2026-08-24T16:29:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetRolloutStuck", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="arobit-forwarder", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="arobit", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: DaemonSet arobit/arobit-forwarder has not finished or progressed for at least 15 minutes. Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [svc] KubeDaemonSetMisScheduled firedfull failure pattern: alert [svc] KubeDaemonSetMisScheduled fired Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | alert | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | Show trend detailsAug 18: 15 · Aug 19: 9 · Aug 20: 9 · Aug 21: 10 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)alert fired 4 time(s) Firing 1: State: Resolved Started: 2026-08-24T16:12:09Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetMisScheduled", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="node-agent", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="velero", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: 1 Pods of DaemonSet velero/node-agent are running where they are not supposed to run. Firing 2: State: Resolved Started: 2026-08-24T16:12:09Z Ended: 2026-08-24T16:24:07Z Severity: Sev3 Labels: alertname="KubeDaemonSetMisScheduled", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="arobit-forwarder", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="arobit", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: 1 Pods of DaemonSet arobit/arobit-forwarder are running where they are not supposed to run. Firing 3: State: Resolved Started: 2026-08-24T16:13:10Z Ended: 2026-08-24T16:23:08Z Severity: Sev3 Labels: alertname="KubeDaemonSetMisScheduled", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="arobit-forwarder", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="arobit", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: 1 Pods of DaemonSet arobit/arobit-forwarder are running where they are not supposed to run. Firing 4: State: Resolved Started: 2026-08-24T16:13:11Z Ended: 2026-08-24T16:23:08Z Severity: Sev3 Labels: alertname="KubeDaemonSetMisScheduled", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", daemonset="node-agent", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="velero", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: 1 Pods of DaemonSet velero/node-agent are running where they are not supposed to run. Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [svc] KubeNodeUnreachable firedfull failure pattern: alert [svc] KubeNodeUnreachable fired Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | alert | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | Show trend detailsAug 18: 7 · Aug 19: 4 · Aug 20: 7 · Aug 21: 9 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)alert fired 2 time(s) Firing 1: State: Resolved Started: 2026-08-24T16:12:59Z Ended: 2026-08-24T16:22:56Z Severity: Sev3 Labels: alertname="KubeNodeUnreachable", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", effect="NoSchedule", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", key="node.kubernetes.io/unreachable", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="prometheus", node="aks-userswft2-20578738-vmss000002", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: aks-userswft2-20578738-vmss000002 is unreachable and some workloads may be rescheduled. Firing 2: State: Resolved Started: 2026-08-24T16:12:59Z Ended: 2026-08-24T16:23:55Z Severity: Sev3 Labels: alertname="KubeNodeUnreachable", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", effect="NoSchedule", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", key="node.kubernetes.io/unreachable", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="prometheus", node="aks-userswft2-20578738-vmss000002", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-wgnw8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: aks-userswft2-20578738-vmss000002 is unreachable and some workloads may be rescheduled. Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [hcp] KubeDeploymentReplicasMismatch firedfull failure pattern: alert [hcp] KubeDeploymentReplicasMismatch fired Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 6 days | alert | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 6 days | Show trend detailsAug 18: 16 · Aug 19: 6 · Aug 20: 8 · Aug 21: 5 · Aug 22: 3 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)alert fired 1 time(s) Firing 1: State: Resolved Started: 2026-08-24T18:15:21Z Ended: 2026-08-24T18:36:22Z Severity: Sev3 Labels: alertname="KubeDeploymentReplicasMismatch", cluster="ci01-j0375808-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", deployment="capi-provider", endpoint="http", environment="ci01", instance="10.128.64.140:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0375808/providers/microsoft.monitor/accounts/hcps-j0375808", namespace="ocm-arohcpci01-2sd37mkptl2qprba4ejkvlbldgl77u9k-r3i6h6d2n5v5b4i", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-hf5qz", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: Deployment ocm-arohcpci01-2sd37mkptl2qprba4ejkvlbldgl77u9k-r3i6h6d2n5v5b4i/capi-provider has not matched the expected number of replicas for longer than 30 minutes. Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [svc] KubeDeploymentReplicasMismatch firedfull failure pattern: alert [svc] KubeDeploymentReplicasMismatch fired Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | alert | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | Show trend detailsAug 18: 1 · Aug 19: 3 · Aug 20: 0 · Aug 21: 1 · Aug 22: 1 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)alert fired 6 time(s) Firing 1: State: Resolved Started: 2026-08-24T18:15:41Z Ended: 2026-08-24T18:21:45Z Severity: Sev3 Labels: alertname="KubeDeploymentReplicasMismatch", cluster="ci01-j0375808-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", deployment="klusterlet-addon-workmgr", endpoint="http", environment="ci01", instance="10.128.64.140:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0375808/providers/microsoft.monitor/accounts/services-j0375808", namespace="klusterlet-2sd37mkptl2qprba4ejkvlbldgl77u9k", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-hf5qz", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: Deployment klusterlet-2sd37mkptl2qprba4ejkvlbldgl77u9k/klusterlet-addon-workmgr has not matched the expected number of replicas for longer than 30 minutes. Firing 2: State: Resolved Started: 2026-08-24T18:15:41Z Ended: 2026-08-24T18:22:41Z Severity: Sev3 Labels: alertname="KubeDeploymentReplicasMismatch", cluster="ci01-j0375808-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", deployment="klusterlet-addon-workmgr", endpoint="http", environment="ci01", instance="10.128.64.140:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0375808/providers/microsoft.monitor/accounts/services-j0375808", namespace="klusterlet-2sd37mkptl2qprba4ejkvlbldgl77u9k", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-hf5qz", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: Deployment klusterlet-2sd37mkptl2qprba4ejkvlbldgl77u9k/klusterlet-addon-workmgr has not matched the expected number of replicas for longer than 30 minutes. Firing 3: State: Resolved Started: 2026-08-24T18:15:41Z Ended: 2026-08-24T18:22:41Z Severity: Sev3 Labels: alertname="KubeDeploymentReplicasMismatch", cluster="ci01-j0375808-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", deployment="governance-policy-framework", endpoint="http", environment="ci01", instance="10.128.64.140:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0375808/providers/microsoft.monitor/accounts/services-j0375808", namespace="klusterlet-2sd37mkptl2qprba4ejkvlbldgl77u9k", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-hf5qz", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: Deployment klusterlet-2sd37mkptl2qprba4ejkvlbldgl77u9k/governance-policy-framework has not matched the expected number of replicas for longer than 30 minutes. Firing 4: State: Resolved Started: 2026-08-24T18:15:41Z Ended: 2026-08-24T18:23:41Z Severity: Sev3 Labels: alertname="KubeDeploymentReplicasMismatch", cluster="ci01-j0375808-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", deployment="governance-policy-framework", endpoint="http", environment="ci01", instance="10.128.64.140:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0375808/providers/microsoft.monitor/accounts/services-j0375808", namespace="klusterlet-2sd37mkptl2qprba4ejkvlbldgl77u9k", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-hf5qz", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: Deployment klusterlet-2sd37mkptl2qprba4ejkvlbldgl77u9k/governance-policy-framework has not matched the expected number of replicas for longer than 30 minutes. Firing 5: State: Resolved Started: 2026-08-24T18:15:41Z Ended: 2026-08-24T18:22:41Z Severity: Sev3 Labels: alertname="KubeDeploymentReplicasMismatch", cluster="ci01-j0375808-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", deployment="config-policy-controller", endpoint="http", environment="ci01", instance="10.128.64.140:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0375808/providers/microsoft.monitor/accounts/services-j0375808", namespace="klusterlet-2sd37mkptl2qprba4ejkvlbldgl77u9k", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-hf5qz", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: Deployment klusterlet-2sd37mkptl2qprba4ejkvlbldgl77u9k/config-policy-controller has not matched the expected number of replicas for longer than 30 minutes. Firing 6: State: Resolved Started: 2026-08-24T18:15:41Z Ended: 2026-08-24T18:23:41Z Severity: Sev3 Labels: alertname="KubeDeploymentReplicasMismatch", cluster="ci01-j0375808-mgmt-2", component="kubernetes-infrastructure", container="kube-state-metrics", deployment="config-policy-controller", endpoint="http", environment="ci01", instance="10.128.64.140:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0375808/providers/microsoft.monitor/accounts/services-j0375808", namespace="klusterlet-2sd37mkptl2qprba4ejkvlbldgl77u9k", pod="arohcp-monitor-kube-state-metrics-5b9fdd8b5b-hf5qz", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning" Description: Deployment klusterlet-2sd37mkptl2qprba4ejkvlbldgl77u9k/config-policy-controller has not matched the expected number of replicas for longer than 30 minutes. Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message cluster deletion did not complete before the deadline; [clusterServiceDeletion] <no_message>; [clusterServiceStatus] <no_message>; ...full failure pattern: ERROR CODE: InternalServerError; detail message cluster deletion did not complete before the deadline; [clusterServiceDeletion] <no_message>; [clusterServiceStatus] <no_message>; [descendantResources] <no_message>; [hostedCluster] <no_message>; provider Microsoft.RedHatOpenShift Signal: FlakeSignal: Flake — present in 3 prior week(s); spread across 5 days | e2e | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 3 prior week(s); spread across 5 days | Show trend detailsAug 18: 1 · Aug 19: 1 · Aug 20: 2 · Aug 21: 2 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | STG | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/util/framework/per_test_framework.go:293]: Unexpected error:
<*errors.joinError | 0xc001c982a0>:
failed to cleanup resource group: at least one hcp cluster failed to delete: failed waiting for hcpCluster="cluster-hs-pre-gdnhnw" in resourcegroup="rg-hs-pre--hm7prkdxbkw6" to finish deleting: GET https://rp.j9012096.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/67ed1acc-0077-4b07-848d-484f46170f44
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/67ed1acc-0077-4b07-848d-484f46170f44",
"name": "67ed1acc-0077-4b07-848d-484f46170f44",
"status": "Failed",
"startTime": "2026-08-24T15:21:50.42303027Z",
"endTime": "2026-08-24T15:45:59.872804145Z",
"error": {
"code": "InternalServerError",
"message": "cluster deletion did not complete before the deadline; [clusterServiceDeletion] \u003cno_message\u003e; [clusterServiceStatus] \u003cno_message\u003e; [descendantResources] \u003cno_message\u003e; [hostedCluster] \u003cno_message\u003e"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/util/framework/per_test_framework.go:293]: Unexpected error:
<*errors.joinError | 0xc001c982a0>:
failed to cleanup resource group: at least one hcp cluster failed to delete: failed waiting for hcpCluster="cluster-hs-pre-gdnhnw" in resourcegroup="rg-hs-pre--hm7prkdxbkw6" to finish deleting: GET https://rp.j9012096.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/67ed1acc-0077-4b07-848d-484f46170f44
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/67ed1acc-0077-4b07-848d-484f46170f44",
"name": "67ed1acc-0077-4b07-848d-484f46170f44",
"status": "Failed",
"startTime": "2026-08-24T15:21:50.42303027Z",
"endTime": "2026-08-24T15:45:59.872804145Z",
"error": {
"code": "InternalServerError",
"message": "cluster deletion did not complete before the deadline; [clusterServiceDeletion] \u003cno_message\u003e; [clusterServiceStatus] \u003cno_message\u003e; [descendantResources] \u003cno_message\u003e; [hostedCluster] \u003cno_message\u003e"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [hcp] KubePodCrashLooping firedfull failure pattern: alert [hcp] KubePodCrashLooping fired Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | alert | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | Show trend detailsAug 18: 2 · Aug 19: 2 · Aug 20: 2 · Aug 21: 0 · Aug 22: 2 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)alert fired 2 time(s) Firing 1: State: Resolved Started: 2026-08-24T18:02:20Z Ended: 2026-08-24T18:38:21Z Severity: Sev3 Labels: alertname="KubePodCrashLooping", cluster="ci01-j0375808-mgmt-2", component="kubernetes-infrastructure", container="etcd", endpoint="http", environment="ci01", instance="10.128.64.140:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0375808/providers/microsoft.monitor/accounts/hcps-j0375808", namespace="ocm-arohcpci01-2sd37mkptl2qprba4ejkvlbldgl77u9k-r3i6h6d2n5v5b4i", pod="etcd-1", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", reason="CrashLoopBackOff", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning", uid="4cebd923-31c4-4c0a-b568-63bf1a281b9a" Description: Pod ocm-arohcpci01-2sd37mkptl2qprba4ejkvlbldgl77u9k-r3i6h6d2n5v5b4i/etcd-1 (etcd) is in waiting state (reason: "CrashLoopBackOff"). Firing 2: State: Resolved Started: 2026-08-24T18:03:20Z Ended: 2026-08-24T18:38:21Z Severity: Sev3 Labels: alertname="KubePodCrashLooping", cluster="ci01-j0375808-mgmt-2", component="kubernetes-infrastructure", container="etcd", endpoint="http", environment="ci01", instance="10.128.64.140:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j0375808/providers/microsoft.monitor/accounts/hcps-j0375808", namespace="ocm-arohcpci01-2sd37mkptl2qprba4ejkvlbldgl77u9k-r3i6h6d2n5v5b4i", pod="etcd-1", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", reason="CrashLoopBackOff", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning", uid="4cebd923-31c4-4c0a-b568-63bf1a281b9a" Description: Pod ocm-arohcpci01-2sd37mkptl2qprba4ejkvlbldgl77u9k-r3i6h6d2n5v5b4i/etcd-1 (etcd) is in waiting state (reason: "CrashLoopBackOff"). Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [svc] BackendControllerQueueDepthHigh firedfull failure pattern: alert [svc] BackendControllerQueueDepthHigh fired Signal: FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | alert | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 4 prior week(s); spread across 5 days | Show trend detailsAug 18: 1 · Aug 19: 3 · Aug 20: 2 · Aug 21: 2 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)alert fired 1 time(s) Firing 1: State: Resolved Started: 2026-08-24T14:40:10Z Ended: 2026-08-24T14:49:11Z Severity: Sev3 Labels: alertname="BackendControllerQueueDepthHigh", cluster="ci01-j6438784-svc", component="backend", name="csstatedump", severity="warning" Description: Backend controller workqueue csstatedump has had a depth > 10 for more than 15 minutes, indicating work is accumulating faster than it can be processed. Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message cluster deletion did not complete before the deadline; [clusterServiceDeletion] ClusterService cluster still exists; [descendantRes...full failure pattern: ERROR CODE: InternalServerError; detail message cluster deletion did not complete before the deadline; [clusterServiceDeletion] ClusterService cluster still exists; [descendantResources] remaining resources: serviceProviderClusters; provider Microsoft.RedHatOpenShift Signal: FlakeSignal: Flake — present in 3 prior week(s); spread across 5 days | e2e | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 3 prior week(s); spread across 5 days | Show trend detailsAug 18: 1 · Aug 19: 1 · Aug 20: 1 · Aug 21: 0 · Aug 22: 1 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/util/framework/per_test_framework.go:293]: Unexpected error:
<*errors.joinError | 0xc0010deff0>:
failed to cleanup resource group: at least one hcp cluster failed to delete: failed waiting for hcpCluster="admin-cred-lifecycle-4cmmkc" in resourcegroup="admin-credential-lifecycle-test-mk5bt6ghv29k" to finish deleting: GET https://rp.j9012096.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/0644f278-df05-4603-a10d-3a66a1f851d7
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/0644f278-df05-4603-a10d-3a66a1f851d7",
"name": "0644f278-df05-4603-a10d-3a66a1f851d7",
"status": "Failed",
"startTime": "2026-08-24T15:17:03.672410989Z",
"endTime": "2026-08-24T15:41:10.046088761Z",
"error": {
"code": "InternalServerError",
"message": "cluster deletion did not complete before the deadline; [clusterServiceDeletion] ClusterService cluster /api/aro_hcp/v1alpha1/clusters/2sd0tjfmep0gog9fd4pfkr5bnpc0n642 still exists (deletion dispatched at 2026-08-24T15:17:04Z); [descendantResources] remaining resources: 1 microsoft.redhatopenshift/hcpopenshiftclusters/serviceProviderClusters"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/util/framework/per_test_framework.go:293]: Unexpected error:
<*errors.joinError | 0xc0010deff0>:
failed to cleanup resource group: at least one hcp cluster failed to delete: failed waiting for hcpCluster="admin-cred-lifecycle-4cmmkc" in resourcegroup="admin-credential-lifecycle-test-mk5bt6ghv29k" to finish deleting: GET https://rp.j9012096.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/0644f278-df05-4603-a10d-3a66a1f851d7
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/0644f278-df05-4603-a10d-3a66a1f851d7",
"name": "0644f278-df05-4603-a10d-3a66a1f851d7",
"status": "Failed",
"startTime": "2026-08-24T15:17:03.672410989Z",
"endTime": "2026-08-24T15:41:10.046088761Z",
"error": {
"code": "InternalServerError",
"message": "cluster deletion did not complete before the deadline; [clusterServiceDeletion] ClusterService cluster /api/aro_hcp/v1alpha1/clusters/2sd0tjfmep0gog9fd4pfkr5bnpc0n642 still exists (deletion dispatched at 2026-08-24T15:17:04Z); [descendantResources] remaining resources: 1 microsoft.redhatopenshift/hcpopenshiftclusters/serviceProviderClusters"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable; host...full failure pattern: ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable; hosted cluster degraded: UnavailableReplicas: [kube-apiserver]; provider Microsoft.RedHatOpenShift Signal: FlakeSignal: Flake — present in 2 prior week(s); spread across 3 days | e2e | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 2 prior week(s); spread across 3 days | Show trend detailsAug 18: 1 · Aug 19: 0 · Aug 20: 1 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_private_ingress.go:103]: failed to create HCP cluster "private-ingress" with private ingress
Unexpected error:
<*fmt.wrapError | 0xc001580bc0>:
failed to create HCP cluster private-ingress: failed waiting for cluster="private-ingress" in resourcegroup="private-ingress-cbbn9kpl7fxv" to finish creating: GET https://rp.j1362816.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/c6f112da-48e3-4cbb-a1b9-a47159e85ff5
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/c6f112da-48e3-4cbb-a1b9-a47159e85ff5",
"name": "c6f112da-48e3-4cbb-a1b9-a47159e85ff5",
"status": "Failed",
"startTime": "2026-08-24T15:18:01.504032647Z",
"endTime": "2026-08-24T15:37:08.017891025Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable: Waiting for components to be available: cluster-policy-controller, azure-cloud-controller-manager, cluster-node-tuning-operator, olm-operator, olm-collect-profiles, cluster-autoscaler, control-plane-pki-operator, kube-controller-manager, community-operators-catalog, redhat-operators-catalog, machine-approver, openshift-controller-manager, cluster-storage-operator, dns-operator, redhat-marketplace-catalog, cluster-version-operator, openshift-route-controller-manager, ingress-operator, kube-scheduler, cluster-image-registry-operator, konnectivity-agent, catalog-operator, cluster-network-operator, ignition-server, ignition-server-proxy, openshift-apiserver, hosted-cluster-config-operator, csi-snapshot-controller-operator, packageserver, certified-operators-catalog; hosted cluster degraded: UnavailableReplicas: kube-apiserver deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_private_ingress.go:103]: failed to create HCP cluster "private-ingress" with private ingress
Unexpected error:
<*fmt.wrapError | 0xc001580bc0>:
failed to create HCP cluster private-ingress: failed waiting for cluster="private-ingress" in resourcegroup="private-ingress-cbbn9kpl7fxv" to finish creating: GET https://rp.j1362816.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/c6f112da-48e3-4cbb-a1b9-a47159e85ff5
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/c6f112da-48e3-4cbb-a1b9-a47159e85ff5",
"name": "c6f112da-48e3-4cbb-a1b9-a47159e85ff5",
"status": "Failed",
"startTime": "2026-08-24T15:18:01.504032647Z",
"endTime": "2026-08-24T15:37:08.017891025Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable: Waiting for components to be available: cluster-policy-controller, azure-cloud-controller-manager, cluster-node-tuning-operator, olm-operator, olm-collect-profiles, cluster-autoscaler, control-plane-pki-operator, kube-controller-manager, community-operators-catalog, redhat-operators-catalog, machine-approver, openshift-controller-manager, cluster-storage-operator, dns-operator, redhat-marketplace-catalog, cluster-version-operator, openshift-route-controller-manager, ingress-operator, kube-scheduler, cluster-image-registry-operator, konnectivity-agent, catalog-operator, cluster-network-operator, ignition-server, ignition-server-proxy, openshift-apiserver, hosted-cluster-config-operator, csi-snapshot-controller-operator, packageserver, certified-operators-catalog; hosted cluster degraded: UnavailableReplicas: kube-apiserver deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
expected 3 nodes to have original label key1=value1full failure pattern: expected 3 nodes to have original label key1=value1 Signal: FlakeSignal: Flake — present in 3 prior week(s); spread across 4 days | e2e | 1 | 0.97%1 of 103 job runs affected | FlakeSignal: Flake — present in 3 prior week(s); spread across 4 days | Show trend detailsAug 18: 2 · Aug 19: 1 · Aug 20: 0 · Aug 21: 2 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/e2e/nodepool_labels_taints.go:259]: expected 3 nodes to have original label key1=value1
Expected
<bool>: ...
to be true
fail [github.com/Azure/ARO-HCP/test/e2e/nodepool_labels_taints.go:259]: expected 3 nodes to have original label key1=value1
Expected
<bool>: ...
to be trueContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
query-basic: command failed: exit status 1full failure pattern: query-basic: command failed: exit status 1 Signal: IndeterminateSignal: Indeterminate — no prior history; active 1 day(s) — New failure pattern — no prior history | e2e | 35 | 33.98%35 of 103 job runs affected | IndeterminateSignal: Indeterminate — no prior history; active 1 day(s) — New failure pattern — no prior history | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 35 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (3)fail [github.com/Azure/ARO-HCP/test/e2e/kusto_logs_present.go:100]: Timed out after 600.001s.
Expected success, but got an error:
<*errors.errorString | 0xc000b66310>:
must-gather CLI test failures:
query-basic: command failed: exit status 1
output: �[37m[00:48:59.861]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:48:54.9335446Z);\nlet _endTime = datetime(2026-08-25T00:48:54.9335452Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-dh6t5tfmwl78'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:48:59.8208131Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;e378a004-3ace-4448-8d86-93b49cc20b45\",\r\n \"activityId\": \"3f29cef9-55f3-4e39-a44c-dc2745aeb329\",\r\n \"subActivityId\": \"9dc2e9b4-43ec-49f2-b454-4f12f219c5db\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"9b9edaeb-e277-470d-a49a-ba5e0f233409\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;e378a004-3ace-4448-8d86-93b49cc20b45 ARID=3f29cef9-55f3-4e39-a44c-dc2745aeb329 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/a5037b89-f047-4fed-bfb0-6c9466286aa0 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/a5037b89-f047-4fed-bfb0-6c9466286aa0 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/9b9edaeb-e277-470d-a49a-ba5e0f233409 \u003e DN.FE.ExecuteQuery/9dc2e9b4-43ec-49f2-b454-4f12f219c5db)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:48:59.984]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:48:54.9335446Z);\nlet _endTime = datetime(2026-08-25T00:48:54.9335452Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-dh6t5tfmwl78'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:48:59.8208131Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;e378a004-3ace-4448-8d86-93b49cc20b45\",\r\n \"activityId\": \"3f29cef9-55f3-4e39-a44c-dc2745aeb329\",\r\n \"subActivityId\": \"9dc2e9b4-43ec-49f2-b454-4f12f219c5db\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"9b9edaeb-e277-470d-a49a-ba5e0f233409\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;e378a004-3ace-4448-8d86-93b49cc20b45 ARID=3f29cef9-55f3-4e39-a44c-dc2745aeb329 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/a5037b89-f047-4fed-bfb0-6c9466286aa0 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/a5037b89-f047-4fed-bfb0-6c9466286aa0 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/9b9edaeb-e277-470d-a49a-ba5e0f233409 \u003e DN.FE.ExecuteQuery/9dc2e9b4-43ec-49f2-b454-4f12f219c5db)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
query-skip-hcp-logs: command failed: exit status 1
output: �[37m[00:49:04.412]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:49:00.0668173Z);\nlet _endTime = datetime(2026-08-25T00:49:00.0668177Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-dh6t5tfmwl78'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:49:04.3919639Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;4c4f5480-a590-4f9a-a459-65dc9658f8ef\",\r\n \"activityId\": \"590d8d9f-c432-4b6c-a57a-14f6cfbdb0ea\",\r\n \"subActivityId\": \"f16f10a7-70e1-48c7-9e9d-209207cfc1e8\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"69739f9c-88ef-46ee-93e1-65af7b0a2416\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;4c4f5480-a590-4f9a-a459-65dc9658f8ef ARID=590d8d9f-c432-4b6c-a57a-14f6cfbdb0ea \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/b66f5a6d-cdfd-418f-82f1-af8b0b7b5e03 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/b66f5a6d-cdfd-418f-82f1-af8b0b7b5e03 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/69739f9c-88ef-46ee-93e1-65af7b0a2416 \u003e DN.FE.ExecuteQuery/f16f10a7-70e1-48c7-9e9d-209207cfc1e8)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:49:04.464]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:49:00.0668173Z);\nlet _endTime = datetime(2026-08-25T00:49:00.0668177Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-dh6t5tfmwl78'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:49:04.3919639Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;4c4f5480-a590-4f9a-a459-65dc9658f8ef\",\r\n \"activityId\": \"590d8d9f-c432-4b6c-a57a-14f6cfbdb0ea\",\r\n \"subActivityId\": \"f16f10a7-70e1-48c7-9e9d-209207cfc1e8\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"69739f9c-88ef-46ee-93e1-65af7b0a2416\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;4c4f5480-a590-4f9a-a459-65dc9658f8ef ARID=590d8d9f-c432-4b6c-a57a-14f6cfbdb0ea \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/b66f5a6d-cdfd-418f-82f1-af8b0b7b5e03 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/b66f5a6d-cdfd-418f-82f1-af8b0b7b5e03 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/69739f9c-88ef-46ee-93e1-65af7b0a2416 \u003e DN.FE.ExecuteQuery/f16f10a7-70e1-48c7-9e9d-209207cfc1e8)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
query-collect-systemd-logs: command failed: exit status 1
output: �[37m[00:49:12.903]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:49:08.6625258Z);\nlet _endTime = datetime(2026-08-25T00:49:08.6625262Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-dh6t5tfmwl78'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:49:12.8815919Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;a5c94555-e249-4511-a807-b8477933031f\",\r\n \"activityId\": \"3a1a3ba5-3fc9-4a7a-bbc2-c2b684f8ab2c\",\r\n \"subActivityId\": \"af4bd91d-24d7-4768-a750-045fb74a6811\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"833ff0b4-2415-4cdd-8b37-d9d4e7e62eaa\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;a5c94555-e249-4511-a807-b8477933031f ARID=3a1a3ba5-3fc9-4a7a-bbc2-c2b684f8ab2c \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/58e55353-9773-463e-ab07-afa8f479d884 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/58e55353-9773-463e-ab07-afa8f479d884 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/833ff0b4-2415-4cdd-8b37-d9d4e7e62eaa \u003e DN.FE.ExecuteQuery/af4bd91d-24d7-4768-a750-045fb74a6811)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:49:13.008]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:49:08.6625258Z);\nlet _endTime = datetime(2026-08-25T00:49:08.6625262Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-dh6t5tfmwl78'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:49:12.8815919Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;a5c94555-e249-4511-a807-b8477933031f\",\r\n \"activityId\": \"3a1a3ba5-3fc9-4a7a-bbc2-c2b684f8ab2c\",\r\n \"subActivityId\": \"af4bd91d-24d7-4768-a750-045fb74a6811\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"833ff0b4-2415-4cdd-8b37-d9d4e7e62eaa\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;a5c94555-e249-4511-a807-b8477933031f ARID=3a1a3ba5-3fc9-4a7a-bbc2-c2b684f8ab2c \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/58e55353-9773-463e-ab07-afa8f479d884 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/58e55353-9773-463e-ab07-afa8f479d884 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/833ff0b4-2415-4cdd-8b37-d9d4e7e62eaa \u003e DN.FE.ExecuteQuery/af4bd91d-24d7-4768-a750-045fb74a6811)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
...
fail [github.com/Azure/ARO-HCP/test/e2e/kusto_logs_present.go:100]: Timed out after 600.001s.
Expected success, but got an error:
<*errors.errorString | 0xc000b66310>:
must-gather CLI test failures:
query-basic: command failed: exit status 1
output: �[37m[00:48:59.861]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:48:54.9335446Z);\nlet _endTime = datetime(2026-08-25T00:48:54.9335452Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-dh6t5tfmwl78'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:48:59.8208131Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;e378a004-3ace-4448-8d86-93b49cc20b45\",\r\n \"activityId\": \"3f29cef9-55f3-4e39-a44c-dc2745aeb329\",\r\n \"subActivityId\": \"9dc2e9b4-43ec-49f2-b454-4f12f219c5db\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"9b9edaeb-e277-470d-a49a-ba5e0f233409\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;e378a004-3ace-4448-8d86-93b49cc20b45 ARID=3f29cef9-55f3-4e39-a44c-dc2745aeb329 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/a5037b89-f047-4fed-bfb0-6c9466286aa0 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/a5037b89-f047-4fed-bfb0-6c9466286aa0 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/9b9edaeb-e277-470d-a49a-ba5e0f233409 \u003e DN.FE.ExecuteQuery/9dc2e9b4-43ec-49f2-b454-4f12f219c5db)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:48:59.984]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:48:54.9335446Z);\nlet _endTime = datetime(2026-08-25T00:48:54.9335452Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-dh6t5tfmwl78'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:48:59.8208131Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;e378a004-3ace-4448-8d86-93b49cc20b45\",\r\n \"activityId\": \"3f29cef9-55f3-4e39-a44c-dc2745aeb329\",\r\n \"subActivityId\": \"9dc2e9b4-43ec-49f2-b454-4f12f219c5db\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"9b9edaeb-e277-470d-a49a-ba5e0f233409\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;e378a004-3ace-4448-8d86-93b49cc20b45 ARID=3f29cef9-55f3-4e39-a44c-dc2745aeb329 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/a5037b89-f047-4fed-bfb0-6c9466286aa0 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/a5037b89-f047-4fed-bfb0-6c9466286aa0 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/9b9edaeb-e277-470d-a49a-ba5e0f233409 \u003e DN.FE.ExecuteQuery/9dc2e9b4-43ec-49f2-b454-4f12f219c5db)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
query-skip-hcp-logs: command failed: exit status 1
output: �[37m[00:49:04.412]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:49:00.0668173Z);\nlet _endTime = datetime(2026-08-25T00:49:00.0668177Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-dh6t5tfmwl78'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:49:04.3919639Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;4c4f5480-a590-4f9a-a459-65dc9658f8ef\",\r\n \"activityId\": \"590d8d9f-c432-4b6c-a57a-14f6cfbdb0ea\",\r\n \"subActivityId\": \"f16f10a7-70e1-48c7-9e9d-209207cfc1e8\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"69739f9c-88ef-46ee-93e1-65af7b0a2416\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;4c4f5480-a590-4f9a-a459-65dc9658f8ef ARID=590d8d9f-c432-4b6c-a57a-14f6cfbdb0ea \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/b66f5a6d-cdfd-418f-82f1-af8b0b7b5e03 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/b66f5a6d-cdfd-418f-82f1-af8b0b7b5e03 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/69739f9c-88ef-46ee-93e1-65af7b0a2416 \u003e DN.FE.ExecuteQuery/f16f10a7-70e1-48c7-9e9d-209207cfc1e8)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:49:04.464]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:49:00.0668173Z);\nlet _endTime = datetime(2026-08-25T00:49:00.0668177Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-dh6t5tfmwl78'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:49:04.3919639Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;4c4f5480-a590-4f9a-a459-65dc9658f8ef\",\r\n \"activityId\": \"590d8d9f-c432-4b6c-a57a-14f6cfbdb0ea\",\r\n \"subActivityId\": \"f16f10a7-70e1-48c7-9e9d-209207cfc1e8\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"69739f9c-88ef-46ee-93e1-65af7b0a2416\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;4c4f5480-a590-4f9a-a459-65dc9658f8ef ARID=590d8d9f-c432-4b6c-a57a-14f6cfbdb0ea \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/b66f5a6d-cdfd-418f-82f1-af8b0b7b5e03 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/b66f5a6d-cdfd-418f-82f1-af8b0b7b5e03 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/69739f9c-88ef-46ee-93e1-65af7b0a2416 \u003e DN.FE.ExecuteQuery/f16f10a7-70e1-48c7-9e9d-209207cfc1e8)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
query-collect-systemd-logs: command failed: exit status 1
output: �[37m[00:49:12.903]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:49:08.6625258Z);\nlet _endTime = datetime(2026-08-25T00:49:08.6625262Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-dh6t5tfmwl78'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:49:12.8815919Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;a5c94555-e249-4511-a807-b8477933031f\",\r\n \"activityId\": \"3a1a3ba5-3fc9-4a7a-bbc2-c2b684f8ab2c\",\r\n \"subActivityId\": \"af4bd91d-24d7-4768-a750-045fb74a6811\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"833ff0b4-2415-4cdd-8b37-d9d4e7e62eaa\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;a5c94555-e249-4511-a807-b8477933031f ARID=3a1a3ba5-3fc9-4a7a-bbc2-c2b684f8ab2c \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/58e55353-9773-463e-ab07-afa8f479d884 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/58e55353-9773-463e-ab07-afa8f479d884 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/833ff0b4-2415-4cdd-8b37-d9d4e7e62eaa \u003e DN.FE.ExecuteQuery/af4bd91d-24d7-4768-a750-045fb74a6811)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:49:13.008]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:49:08.6625258Z);\nlet _endTime = datetime(2026-08-25T00:49:08.6625262Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-dh6t5tfmwl78'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:49:12.8815919Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;a5c94555-e249-4511-a807-b8477933031f\",\r\n \"activityId\": \"3a1a3ba5-3fc9-4a7a-bbc2-c2b684f8ab2c\",\r\n \"subActivityId\": \"af4bd91d-24d7-4768-a750-045fb74a6811\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"833ff0b4-2415-4cdd-8b37-d9d4e7e62eaa\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;a5c94555-e249-4511-a807-b8477933031f ARID=3a1a3ba5-3fc9-4a7a-bbc2-c2b684f8ab2c \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/58e55353-9773-463e-ab07-afa8f479d884 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/58e55353-9773-463e-ab07-afa8f479d884 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/833ff0b4-2415-4cdd-8b37-d9d4e7e62eaa \u003e DN.FE.ExecuteQuery/af4bd91d-24d7-4768-a750-045fb74a6811)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
...fail [github.com/Azure/ARO-HCP/test/e2e/kusto_logs_present.go:100]: Timed out after 600.009s.
Expected success, but got an error:
<*errors.errorString | 0xc00083c410>:
must-gather CLI test failures:
query-basic: command failed: exit status 1
output: �[37m[00:23:58.357]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:23:55.0351655Z);\nlet _endTime = datetime(2026-08-25T00:23:55.0351661Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-b5s7zllhqcpk'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:23:58.3330999Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;125f7c3f-8e3c-4adc-92d3-11f56e27a65b\",\r\n \"activityId\": \"442abb20-5490-45f5-ba70-3bca3bed0639\",\r\n \"subActivityId\": \"b9132e8a-1c55-4f09-8982-f344852f0eee\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"99a04993-dace-423f-a879-2fbeee2a7b84\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;125f7c3f-8e3c-4adc-92d3-11f56e27a65b ARID=442abb20-5490-45f5-ba70-3bca3bed0639 \u003e GW.Http.CallContext/442abb20-5490-45f5-ba70-3bca3bed0639 \u003e GW.ExecuteQuery/99a04993-dace-423f-a879-2fbeee2a7b84 \u003e DN.FE.ExecuteQuery/b9132e8a-1c55-4f09-8982-f344852f0eee)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:23:58.411]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:23:55.0351655Z);\nlet _endTime = datetime(2026-08-25T00:23:55.0351661Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-b5s7zllhqcpk'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:23:58.3330999Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;125f7c3f-8e3c-4adc-92d3-11f56e27a65b\",\r\n \"activityId\": \"442abb20-5490-45f5-ba70-3bca3bed0639\",\r\n \"subActivityId\": \"b9132e8a-1c55-4f09-8982-f344852f0eee\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"99a04993-dace-423f-a879-2fbeee2a7b84\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;125f7c3f-8e3c-4adc-92d3-11f56e27a65b ARID=442abb20-5490-45f5-ba70-3bca3bed0639 \u003e GW.Http.CallContext/442abb20-5490-45f5-ba70-3bca3bed0639 \u003e GW.ExecuteQuery/99a04993-dace-423f-a879-2fbeee2a7b84 \u003e DN.FE.ExecuteQuery/b9132e8a-1c55-4f09-8982-f344852f0eee)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
query-skip-hcp-logs: command failed: exit status 1
output: �[37m[00:24:01.787]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:23:58.4821435Z);\nlet _endTime = datetime(2026-08-25T00:23:58.4821442Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-b5s7zllhqcpk'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:24:01.7618170Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;7425e14f-788f-421a-b19d-fdd4f5d966b5\",\r\n \"activityId\": \"d2adae1e-7609-4016-b166-5a8ff8be55d8\",\r\n \"subActivityId\": \"7c0e685c-e2d3-436f-89bc-9d9fa3141b61\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"8cf560a3-f257-4d9d-900b-aaf46185fe8b\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;7425e14f-788f-421a-b19d-fdd4f5d966b5 ARID=d2adae1e-7609-4016-b166-5a8ff8be55d8 \u003e GW.Http.CallContext/d2adae1e-7609-4016-b166-5a8ff8be55d8 \u003e GW.ExecuteQuery/8cf560a3-f257-4d9d-900b-aaf46185fe8b \u003e DN.FE.ExecuteQuery/7c0e685c-e2d3-436f-89bc-9d9fa3141b61)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:24:01.845]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:23:58.4821435Z);\nlet _endTime = datetime(2026-08-25T00:23:58.4821442Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-b5s7zllhqcpk'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:24:01.7618170Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;7425e14f-788f-421a-b19d-fdd4f5d966b5\",\r\n \"activityId\": \"d2adae1e-7609-4016-b166-5a8ff8be55d8\",\r\n \"subActivityId\": \"7c0e685c-e2d3-436f-89bc-9d9fa3141b61\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"8cf560a3-f257-4d9d-900b-aaf46185fe8b\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;7425e14f-788f-421a-b19d-fdd4f5d966b5 ARID=d2adae1e-7609-4016-b166-5a8ff8be55d8 \u003e GW.Http.CallContext/d2adae1e-7609-4016-b166-5a8ff8be55d8 \u003e GW.ExecuteQuery/8cf560a3-f257-4d9d-900b-aaf46185fe8b \u003e DN.FE.ExecuteQuery/7c0e685c-e2d3-436f-89bc-9d9fa3141b61)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
query-collect-systemd-logs: command failed: exit status 1
output: �[37m[00:24:08.682]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:24:05.4122339Z);\nlet _endTime = datetime(2026-08-25T00:24:05.4122344Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-b5s7zllhqcpk'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:24:08.6621602Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;6e8eb1dd-b3b2-4c08-bb8a-fabaa900a66f\",\r\n \"activityId\": \"ed3e04f4-6a49-4e81-9b00-d4558d9cddcc\",\r\n \"subActivityId\": \"4cad9bba-6029-42b0-a77f-5c14bc4e8a84\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"58dfa59a-bdfb-430f-af23-104aaad94cc9\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;6e8eb1dd-b3b2-4c08-bb8a-fabaa900a66f ARID=ed3e04f4-6a49-4e81-9b00-d4558d9cddcc \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/7d4daa74-783c-457a-8256-bc6951fb2784 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/7d4daa74-783c-457a-8256-bc6951fb2784 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/58dfa59a-bdfb-430f-af23-104aaad94cc9 \u003e DN.FE.ExecuteQuery/4cad9bba-6029-42b0-a77f-5c14bc4e8a84)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:24:08.778]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:24:05.4122339Z);\nlet _endTime = datetime(2026-08-25T00:24:05.4122344Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-b5s7zllhqcpk'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:24:08.6621602Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;6e8eb1dd-b3b2-4c08-bb8a-fabaa900a66f\",\r\n \"activityId\": \"ed3e04f4-6a49-4e81-9b00-d4558d9cddcc\",\r\n \"subActivityId\": \"4cad9bba-6029-42b0-a77f-5c14bc4e8a84\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"58dfa59a-bdfb-430f-af23-104aaad94cc9\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;6e8eb1dd-b3b2-4c08-bb8a-fabaa900a66f ARID=ed3e04f4-6a49-4e81-9b00-d4558d9cddcc \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/7d4daa74-783c-457a-8256-bc6951fb2784 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/7d4daa74-783c-457a-8256-bc6951fb2784 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/58dfa59a-bdfb-430f-af23-104aaad94cc9 \u003e DN.FE.ExecuteQuery/4cad9bba-6029-42b0-a77f-5c14bc4e8a84)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
...
fail [github.com/Azure/ARO-HCP/test/e2e/kusto_logs_present.go:100]: Timed out after 600.009s.
Expected success, but got an error:
<*errors.errorString | 0xc00083c410>:
must-gather CLI test failures:
query-basic: command failed: exit status 1
output: �[37m[00:23:58.357]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:23:55.0351655Z);\nlet _endTime = datetime(2026-08-25T00:23:55.0351661Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-b5s7zllhqcpk'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:23:58.3330999Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;125f7c3f-8e3c-4adc-92d3-11f56e27a65b\",\r\n \"activityId\": \"442abb20-5490-45f5-ba70-3bca3bed0639\",\r\n \"subActivityId\": \"b9132e8a-1c55-4f09-8982-f344852f0eee\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"99a04993-dace-423f-a879-2fbeee2a7b84\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;125f7c3f-8e3c-4adc-92d3-11f56e27a65b ARID=442abb20-5490-45f5-ba70-3bca3bed0639 \u003e GW.Http.CallContext/442abb20-5490-45f5-ba70-3bca3bed0639 \u003e GW.ExecuteQuery/99a04993-dace-423f-a879-2fbeee2a7b84 \u003e DN.FE.ExecuteQuery/b9132e8a-1c55-4f09-8982-f344852f0eee)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:23:58.411]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:23:55.0351655Z);\nlet _endTime = datetime(2026-08-25T00:23:55.0351661Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-b5s7zllhqcpk'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:23:58.3330999Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;125f7c3f-8e3c-4adc-92d3-11f56e27a65b\",\r\n \"activityId\": \"442abb20-5490-45f5-ba70-3bca3bed0639\",\r\n \"subActivityId\": \"b9132e8a-1c55-4f09-8982-f344852f0eee\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"99a04993-dace-423f-a879-2fbeee2a7b84\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;125f7c3f-8e3c-4adc-92d3-11f56e27a65b ARID=442abb20-5490-45f5-ba70-3bca3bed0639 \u003e GW.Http.CallContext/442abb20-5490-45f5-ba70-3bca3bed0639 \u003e GW.ExecuteQuery/99a04993-dace-423f-a879-2fbeee2a7b84 \u003e DN.FE.ExecuteQuery/b9132e8a-1c55-4f09-8982-f344852f0eee)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
query-skip-hcp-logs: command failed: exit status 1
output: �[37m[00:24:01.787]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:23:58.4821435Z);\nlet _endTime = datetime(2026-08-25T00:23:58.4821442Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-b5s7zllhqcpk'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:24:01.7618170Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;7425e14f-788f-421a-b19d-fdd4f5d966b5\",\r\n \"activityId\": \"d2adae1e-7609-4016-b166-5a8ff8be55d8\",\r\n \"subActivityId\": \"7c0e685c-e2d3-436f-89bc-9d9fa3141b61\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"8cf560a3-f257-4d9d-900b-aaf46185fe8b\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;7425e14f-788f-421a-b19d-fdd4f5d966b5 ARID=d2adae1e-7609-4016-b166-5a8ff8be55d8 \u003e GW.Http.CallContext/d2adae1e-7609-4016-b166-5a8ff8be55d8 \u003e GW.ExecuteQuery/8cf560a3-f257-4d9d-900b-aaf46185fe8b \u003e DN.FE.ExecuteQuery/7c0e685c-e2d3-436f-89bc-9d9fa3141b61)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:24:01.845]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:23:58.4821435Z);\nlet _endTime = datetime(2026-08-25T00:23:58.4821442Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-b5s7zllhqcpk'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:24:01.7618170Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;7425e14f-788f-421a-b19d-fdd4f5d966b5\",\r\n \"activityId\": \"d2adae1e-7609-4016-b166-5a8ff8be55d8\",\r\n \"subActivityId\": \"7c0e685c-e2d3-436f-89bc-9d9fa3141b61\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"8cf560a3-f257-4d9d-900b-aaf46185fe8b\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;7425e14f-788f-421a-b19d-fdd4f5d966b5 ARID=d2adae1e-7609-4016-b166-5a8ff8be55d8 \u003e GW.Http.CallContext/d2adae1e-7609-4016-b166-5a8ff8be55d8 \u003e GW.ExecuteQuery/8cf560a3-f257-4d9d-900b-aaf46185fe8b \u003e DN.FE.ExecuteQuery/7c0e685c-e2d3-436f-89bc-9d9fa3141b61)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
query-collect-systemd-logs: command failed: exit status 1
output: �[37m[00:24:08.682]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:24:05.4122339Z);\nlet _endTime = datetime(2026-08-25T00:24:05.4122344Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-b5s7zllhqcpk'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:24:08.6621602Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;6e8eb1dd-b3b2-4c08-bb8a-fabaa900a66f\",\r\n \"activityId\": \"ed3e04f4-6a49-4e81-9b00-d4558d9cddcc\",\r\n \"subActivityId\": \"4cad9bba-6029-42b0-a77f-5c14bc4e8a84\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"58dfa59a-bdfb-430f-af23-104aaad94cc9\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;6e8eb1dd-b3b2-4c08-bb8a-fabaa900a66f ARID=ed3e04f4-6a49-4e81-9b00-d4558d9cddcc \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/7d4daa74-783c-457a-8256-bc6951fb2784 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/7d4daa74-783c-457a-8256-bc6951fb2784 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/58dfa59a-bdfb-430f-af23-104aaad94cc9 \u003e DN.FE.ExecuteQuery/4cad9bba-6029-42b0-a77f-5c14bc4e8a84)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:24:08.778]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:24:05.4122339Z);\nlet _endTime = datetime(2026-08-25T00:24:05.4122344Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-b5s7zllhqcpk'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:24:08.6621602Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;6e8eb1dd-b3b2-4c08-bb8a-fabaa900a66f\",\r\n \"activityId\": \"ed3e04f4-6a49-4e81-9b00-d4558d9cddcc\",\r\n \"subActivityId\": \"4cad9bba-6029-42b0-a77f-5c14bc4e8a84\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"58dfa59a-bdfb-430f-af23-104aaad94cc9\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;6e8eb1dd-b3b2-4c08-bb8a-fabaa900a66f ARID=ed3e04f4-6a49-4e81-9b00-d4558d9cddcc \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/7d4daa74-783c-457a-8256-bc6951fb2784 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/7d4daa74-783c-457a-8256-bc6951fb2784 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/58dfa59a-bdfb-430f-af23-104aaad94cc9 \u003e DN.FE.ExecuteQuery/4cad9bba-6029-42b0-a77f-5c14bc4e8a84)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
...fail [github.com/Azure/ARO-HCP/test/e2e/kusto_logs_present.go:100]: Timed out after 600.003s.
Expected success, but got an error:
<*errors.errorString | 0xc000626a40>:
must-gather CLI test failures:
query-basic: command failed: exit status 1
output: �[37m[00:14:56.798]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:14:53.4068554Z);\nlet _endTime = datetime(2026-08-25T00:14:53.4068559Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-jk7hknvqffmd'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:14:56.7779164Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;cb2874aa-23b9-4e90-82b8-8dbdc8ccb12d\",\r\n \"activityId\": \"697030e7-6169-48d4-afab-e55af5423283\",\r\n \"subActivityId\": \"7025f3d5-21ad-4b66-b1ec-6016c4d8ba2f\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"e5629797-dc74-431c-89f2-26cc2ad935e1\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;cb2874aa-23b9-4e90-82b8-8dbdc8ccb12d ARID=697030e7-6169-48d4-afab-e55af5423283 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/d94d70a1-2ca3-4462-b247-1c225640a5b0 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/d94d70a1-2ca3-4462-b247-1c225640a5b0 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/e5629797-dc74-431c-89f2-26cc2ad935e1 \u003e DN.FE.ExecuteQuery/7025f3d5-21ad-4b66-b1ec-6016c4d8ba2f)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:14:56.846]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:14:53.4068554Z);\nlet _endTime = datetime(2026-08-25T00:14:53.4068559Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-jk7hknvqffmd'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:14:56.7779164Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;cb2874aa-23b9-4e90-82b8-8dbdc8ccb12d\",\r\n \"activityId\": \"697030e7-6169-48d4-afab-e55af5423283\",\r\n \"subActivityId\": \"7025f3d5-21ad-4b66-b1ec-6016c4d8ba2f\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"e5629797-dc74-431c-89f2-26cc2ad935e1\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;cb2874aa-23b9-4e90-82b8-8dbdc8ccb12d ARID=697030e7-6169-48d4-afab-e55af5423283 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/d94d70a1-2ca3-4462-b247-1c225640a5b0 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/d94d70a1-2ca3-4462-b247-1c225640a5b0 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/e5629797-dc74-431c-89f2-26cc2ad935e1 \u003e DN.FE.ExecuteQuery/7025f3d5-21ad-4b66-b1ec-6016c4d8ba2f)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
query-skip-hcp-logs: command failed: exit status 1
output: �[37m[00:15:00.265]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:14:56.9537642Z);\nlet _endTime = datetime(2026-08-25T00:14:56.9537648Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-jk7hknvqffmd'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:15:00.2458657Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;d3045808-57d8-4c23-8f73-affb30f6a130\",\r\n \"activityId\": \"60865211-a5af-4242-8bd8-14b7a1f199ba\",\r\n \"subActivityId\": \"865c926e-620f-4107-9fd8-43ac9ee494b3\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"70a56a06-7a14-4b6c-b645-990ed51e93ae\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;d3045808-57d8-4c23-8f73-affb30f6a130 ARID=60865211-a5af-4242-8bd8-14b7a1f199ba \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/230b95d6-3e5a-4f20-bb04-f8704206f9bc \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/230b95d6-3e5a-4f20-bb04-f8704206f9bc \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/70a56a06-7a14-4b6c-b645-990ed51e93ae \u003e DN.FE.ExecuteQuery/865c926e-620f-4107-9fd8-43ac9ee494b3)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:15:00.327]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:14:56.9537642Z);\nlet _endTime = datetime(2026-08-25T00:14:56.9537648Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-jk7hknvqffmd'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:15:00.2458657Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;d3045808-57d8-4c23-8f73-affb30f6a130\",\r\n \"activityId\": \"60865211-a5af-4242-8bd8-14b7a1f199ba\",\r\n \"subActivityId\": \"865c926e-620f-4107-9fd8-43ac9ee494b3\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"70a56a06-7a14-4b6c-b645-990ed51e93ae\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;d3045808-57d8-4c23-8f73-affb30f6a130 ARID=60865211-a5af-4242-8bd8-14b7a1f199ba \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/230b95d6-3e5a-4f20-bb04-f8704206f9bc \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/230b95d6-3e5a-4f20-bb04-f8704206f9bc \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/70a56a06-7a14-4b6c-b645-990ed51e93ae \u003e DN.FE.ExecuteQuery/865c926e-620f-4107-9fd8-43ac9ee494b3)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
query-collect-systemd-logs: command failed: exit status 1
output: �[37m[00:15:08.703]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:15:04.4342198Z);\nlet _endTime = datetime(2026-08-25T00:15:04.4342235Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-jk7hknvqffmd'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:15:08.6768100Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;2c78ad4b-fe52-4282-9d01-7f8d1f03d199\",\r\n \"activityId\": \"a7bce1a5-15e0-4dd5-9933-19a9ff7b6f19\",\r\n \"subActivityId\": \"dda68eb2-c740-498f-b62e-c74a25b621f7\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"4d4ca23c-dcca-4b88-9136-446102508ccb\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;2c78ad4b-fe52-4282-9d01-7f8d1f03d199 ARID=a7bce1a5-15e0-4dd5-9933-19a9ff7b6f19 \u003e GW.Http.CallContext/a7bce1a5-15e0-4dd5-9933-19a9ff7b6f19 \u003e GW.ExecuteQuery/4d4ca23c-dcca-4b88-9136-446102508ccb \u003e DN.FE.ExecuteQuery/dda68eb2-c740-498f-b62e-c74a25b621f7)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:15:08.705]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "context canceled",
"name": "cosmosResourceSnapshots"
}�[0m
�[37m[00:15:09.096]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:15:04.4342198Z);\nlet _endTime = datetime(2026-08-25T00:15:04.4342235Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-jk7hknvqffmd'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:15:08.6768100Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;2c78ad4b-fe52-4282-9d01-7f8d1f03d199\",\r\n \"activityId\": \"a7bce1a5-15e0-4dd5-9933-19a9ff7b6f19\",\r\n \"subActivityId\": \"dda68eb2-c740-498f-b62e-c74a25b621f7\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"4d4ca23c-dcca-4b88-9136-446102508ccb\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;2c78ad4b-fe52-4282-9d01-7f8d1f03d199 ARID=a7bce1a5-15e0-4dd5-9933-19a9ff7b6f19 \u003e GW.Http.CallContext/a7bce1a5-15e0-4dd5-9933-19a9ff7b6f19 \u003e GW.ExecuteQuery/4d4ca23c-dcca-4b88-9136-446102508ccb \u003e DN.FE.ExecuteQuery/dda68eb2-c740-498f-b62e-c74a25b621f7)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
...
fail [github.com/Azure/ARO-HCP/test/e2e/kusto_logs_present.go:100]: Timed out after 600.003s.
Expected success, but got an error:
<*errors.errorString | 0xc000626a40>:
must-gather CLI test failures:
query-basic: command failed: exit status 1
output: �[37m[00:14:56.798]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:14:53.4068554Z);\nlet _endTime = datetime(2026-08-25T00:14:53.4068559Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-jk7hknvqffmd'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:14:56.7779164Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;cb2874aa-23b9-4e90-82b8-8dbdc8ccb12d\",\r\n \"activityId\": \"697030e7-6169-48d4-afab-e55af5423283\",\r\n \"subActivityId\": \"7025f3d5-21ad-4b66-b1ec-6016c4d8ba2f\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"e5629797-dc74-431c-89f2-26cc2ad935e1\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;cb2874aa-23b9-4e90-82b8-8dbdc8ccb12d ARID=697030e7-6169-48d4-afab-e55af5423283 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/d94d70a1-2ca3-4462-b247-1c225640a5b0 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/d94d70a1-2ca3-4462-b247-1c225640a5b0 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/e5629797-dc74-431c-89f2-26cc2ad935e1 \u003e DN.FE.ExecuteQuery/7025f3d5-21ad-4b66-b1ec-6016c4d8ba2f)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:14:56.846]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:14:53.4068554Z);\nlet _endTime = datetime(2026-08-25T00:14:53.4068559Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-jk7hknvqffmd'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:14:56.7779164Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;cb2874aa-23b9-4e90-82b8-8dbdc8ccb12d\",\r\n \"activityId\": \"697030e7-6169-48d4-afab-e55af5423283\",\r\n \"subActivityId\": \"7025f3d5-21ad-4b66-b1ec-6016c4d8ba2f\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"e5629797-dc74-431c-89f2-26cc2ad935e1\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;cb2874aa-23b9-4e90-82b8-8dbdc8ccb12d ARID=697030e7-6169-48d4-afab-e55af5423283 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/d94d70a1-2ca3-4462-b247-1c225640a5b0 \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/d94d70a1-2ca3-4462-b247-1c225640a5b0 \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/e5629797-dc74-431c-89f2-26cc2ad935e1 \u003e DN.FE.ExecuteQuery/7025f3d5-21ad-4b66-b1ec-6016c4d8ba2f)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
query-skip-hcp-logs: command failed: exit status 1
output: �[37m[00:15:00.265]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:14:56.9537642Z);\nlet _endTime = datetime(2026-08-25T00:14:56.9537648Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-jk7hknvqffmd'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:15:00.2458657Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;d3045808-57d8-4c23-8f73-affb30f6a130\",\r\n \"activityId\": \"60865211-a5af-4242-8bd8-14b7a1f199ba\",\r\n \"subActivityId\": \"865c926e-620f-4107-9fd8-43ac9ee494b3\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"70a56a06-7a14-4b6c-b645-990ed51e93ae\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;d3045808-57d8-4c23-8f73-affb30f6a130 ARID=60865211-a5af-4242-8bd8-14b7a1f199ba \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/230b95d6-3e5a-4f20-bb04-f8704206f9bc \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/230b95d6-3e5a-4f20-bb04-f8704206f9bc \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/70a56a06-7a14-4b6c-b645-990ed51e93ae \u003e DN.FE.ExecuteQuery/865c926e-620f-4107-9fd8-43ac9ee494b3)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:15:00.327]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:14:56.9537642Z);\nlet _endTime = datetime(2026-08-25T00:14:56.9537648Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-jk7hknvqffmd'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:15:00.2458657Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;d3045808-57d8-4c23-8f73-affb30f6a130\",\r\n \"activityId\": \"60865211-a5af-4242-8bd8-14b7a1f199ba\",\r\n \"subActivityId\": \"865c926e-620f-4107-9fd8-43ac9ee494b3\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"70a56a06-7a14-4b6c-b645-990ed51e93ae\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;d3045808-57d8-4c23-8f73-affb30f6a130 ARID=60865211-a5af-4242-8bd8-14b7a1f199ba \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/230b95d6-3e5a-4f20-bb04-f8704206f9bc \u003e KD.Query.Client.ExecuteQueryAsKustoDataStream/230b95d6-3e5a-4f20-bb04-f8704206f9bc \u003e P.Grpc.Service.ExecuteQueryAsKustoDataStream2/70a56a06-7a14-4b6c-b645-990ed51e93ae \u003e DN.FE.ExecuteQuery/865c926e-620f-4107-9fd8-43ac9ee494b3)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
query-collect-systemd-logs: command failed: exit status 1
output: �[37m[00:15:08.703]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:15:04.4342198Z);\nlet _endTime = datetime(2026-08-25T00:15:04.4342235Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-jk7hknvqffmd'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:15:08.6768100Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;2c78ad4b-fe52-4282-9d01-7f8d1f03d199\",\r\n \"activityId\": \"a7bce1a5-15e0-4dd5-9933-19a9ff7b6f19\",\r\n \"subActivityId\": \"dda68eb2-c740-498f-b62e-c74a25b621f7\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"4d4ca23c-dcca-4b88-9136-446102508ccb\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;2c78ad4b-fe52-4282-9d01-7f8d1f03d199 ARID=a7bce1a5-15e0-4dd5-9933-19a9ff7b6f19 \u003e GW.Http.CallContext/a7bce1a5-15e0-4dd5-9933-19a9ff7b6f19 \u003e GW.ExecuteQuery/4d4ca23c-dcca-4b88-9136-446102508ccb \u003e DN.FE.ExecuteQuery/dda68eb2-c740-498f-b62e-c74a25b621f7)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}",
"name": "detailedInfraOrchestrationLogs"
}�[0m
�[37m[00:15:08.705]�[0m �[91mERROR:�[0m �[97mQuery failed�[0m �[90m{
"err": "context canceled",
"name": "cosmosResourceSnapshots"
}�[0m
�[37m[00:15:09.096]�[0m �[91mERROR:�[0m �[97mcommand failed�[0m �[90m{
"err": "failed to execute custom logs query: error during query execution: failed to execute query: failed to execute query: Op(OpQuery): Kind(KHTTPError): error from Kusto endpoint, With query: let _startTime = datetime(2026-08-24T00:15:04.4342198Z);\nlet _endTime = datetime(2026-08-25T00:15:04.4342235Z);\n// Extract cluster ID from CS logs using cid field\nlet cluster_id = toscalar(\n database('ServiceLogs').table('clustersServiceLogs')\n | where timestamp between (_startTime .. _endTime)\n | where resource_id has 'kusto-logs-jk7hknvqffmd'\n | where isnotempty(log.cid)\n | distinct tostring(log.cid)\n | limit 1\n);\n// Find all Maestro resource IDs for this cluster\nlet maestro_resource_ids = (\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | extend resourceId = extract(@'resourceid[=:\"]+([a-f0-9\\-]+)', 1, logStr)\n | where isnotempty(resourceId)\n | distinct resourceId\n);\n// Infrastructure \u0026 Orchestration Logs - Maestro, Hypershift, and ACM Agent\nunion\n (\n // Maestro logs: Resource bundle operations for this cluster and its associated resources\n // Includes: resource creation, updates, deletions, status events, gRPC communications\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"maestro\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id or logStr has_any (maestro_resource_ids)\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // Hypershift logs: Hypershift operator reconciliation of HostedCluster and NodePool CRs\n // Includes the CR reconciliation logs\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"hypershift\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n ),\n (\n // ACM Agent logs: ManifestWork reconciliation on the managed cluster\n // Includes: work agent operations, manifest application, status feedback\n database('ServiceLogs').table('containerLogs')\n | where timestamp between (_startTime .. _endTime)\n | where namespace_name has \"open-cluster-management-agent\"\n | extend logStr = tostring(log)\n | where logStr has cluster_id\n | project timestamp, container_name, msg = logStr, log\n )\n| order by timestamp asc(400 Bad Request):\n{\r\n \"error\": {\r\n \"code\": \"General_BadRequest\",\r\n \"message\": \"Request is invalid and cannot be executed.\",\r\n \"@type\": \"Kusto.Data.Exceptions.RelopSemanticException\",\r\n \"@message\": \"Relop semantic error: SEM0026: The arguments array exceeded the allowed limit (allowed=10,000)\",\r\n \"@failureCode\": 400,\r\n \"@context\": {\r\n \"timestamp\": \"2026-08-25T00:15:08.6768100Z\",\r\n \"serviceAlias\": \"HCP-DEV-US-2\",\r\n \"clientRequestId\": \"KGC.execute;2c78ad4b-fe52-4282-9d01-7f8d1f03d199\",\r\n \"activityId\": \"a7bce1a5-15e0-4dd5-9933-19a9ff7b6f19\",\r\n \"subActivityId\": \"dda68eb2-c740-498f-b62e-c74a25b621f7\",\r\n \"activityType\": \"DN.FE.ExecuteQuery\",\r\n \"parentActivityId\": \"4d4ca23c-dcca-4b88-9136-446102508ccb\",\r\n \"activityStack\": \"(Activity stack: CRID=KGC.execute;2c78ad4b-fe52-4282-9d01-7f8d1f03d199 ARID=a7bce1a5-15e0-4dd5-9933-19a9ff7b6f19 \u003e GW.Http.CallContext/a7bce1a5-15e0-4dd5-9933-19a9ff7b6f19 \u003e GW.ExecuteQuery/4d4ca23c-dcca-4b88-9136-446102508ccb \u003e DN.FE.ExecuteQuery/dda68eb2-c740-498f-b62e-c74a25b621f7)\",\r\n \"serviceFarm\": null\r\n },\r\n \"@permanent\": true\r\n }\r\n}"
}�[0m
...Contributing tests (1)
Affected runs (35) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: DeploymentFailed; detail code OverconstrainedZonalAllocationRequest; detail message Allocation failed. VM(s) with the following constraints cannot be allocated, becaus...full failure pattern: ERROR CODE: DeploymentFailed; detail code OverconstrainedZonalAllocationRequest; detail message Allocation failed. VM(s) with the following constraints cannot be allocated, because the condition is too restrictive. Please remove some constraints and try again. Constraints applied are:\\ - Availability Zone\\ - Differencing (Ephemeral) Disks\\ - Networking Constraints (such as Accelerated Networking or IPv6)\\ - VM Size\\; provider Microsoft.Compute Signal: IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 2 day(s) | provision | 28 | 27.18%28 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 2 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 6 · Aug 22: 0 · Aug 23: 0 · Aug 24: 28 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (3)time=2026-08-24T13:26:56.540Z level=INFO msg="Running step." serviceGroup=Microsoft.Azure.ARO.HCP.Service.Infra resourceGroup=service step=cluster description="Step cluster\n Kind: ARM\n Template: templates/svc-cluster.bicep\n Parameters: configurations/svc-cluster.tmpl.bicepparam"
time=2026-08-24T13:26:58.762Z level=DEBUG+2 msg="Starting ARM deployment" serviceGroup=Microsoft.Azure.ARO.HCP.Service.Infra resourceGroup=service step=cluster
time=2026-08-24T13:27:00.485Z level=DEBUG+3 msg="Deployment started" serviceGroup=Microsoft.Azure.ARO.HCP.Service.Infra resourceGroup=service step=cluster deployment=abeb20b90f6fb4be4d4a69c08ed3bbe8b5be5f3247506ae425e325b85b345969 portal=https://ms.portal.azure.com/#view/Microsoft_Azure_Resources/DeploymentDetails.MenuView/~/overview/id/%2Fsubscriptions%2FXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX%2FresourceGroups%2Fhcp-underlay-ci01-j8084480-svc%2Fproviders%2FMicrosoft.Resources%2Fdeployments%2Fabeb20b90f6fb4be4d4a69c08ed3bbe8b5be5f3247506ae425e325b85b345969
time=2026-08-24T13:35:32.482Z level=ERROR msg="Step errored." serviceGroup=Microsoft.Azure.ARO.HCP.Service.Infra resourceGroup=service step=cluster err="failed to run ARM step: failed to poll deployment: failed to wait for deployment completion: GET https://management.azure.com/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j8084480-svc/providers/Microsoft.Resources/deployments/abeb20b90f6fb4be4d4a69c08ed3bbe8b5be5f3247506ae425e325b85b345969/operationStatuses/08584140288659502671\n--------------------------------------------------------------------------------\nRESPONSE 200: 200 OK\nERROR CODE: DeploymentFailed\n--------------------------------------------------------------------------------\n{\n \"status\": \"Failed\",\n \"error\": {\n \"code\": \"DeploymentFailed\",\n \"message\": \"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\",\n \"details\": [\n {\n \"code\": \"Conflict\",\n \"message\": \"{\\r\\n \\\"status\\\": \\\"Failed\\\",\\r\\n \\\"error\\\": {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"DeploymentFailed\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j8084480-svc/providers/Microsoft.Resources/deployments/cluster-kyv757fphd2fm\\\",\\r\\n \\\"message\\\": \\\"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j8084480-svc/providers/Microsoft.Resources/deployments/user-agent-pools\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"DeploymentFailed\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j8084480-svc/providers/Microsoft.Resources/deployments/user-agent-pools\\\",\\r\\n \\\"message\\\": \\\"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j8084480-svc/providers/Microsoft.ContainerService/managedClusters/ci01-j8084480-svc/agentPools/u64d8dsv61\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"OverconstrainedZonalAllocationRequest\\\",\\r\\n \\\"message\\\": \\\"Allocation failed on resource '/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j8084480-svc-aks1/providers/Microsoft.Compute/virtualMachineScaleSets/aks-u64d8dsv61-24652102-vmss'. Please try alternative sizes 'Standard_D8s_v6' (zones 3, 2), 'Standard_E8-4s_v6' (zones 3, 2), 'Standard_E8-4ds_v6' (zones 2, 3) instead for higher likelihood of success. Read more about improving likelihood of operation success at https://aka.ms/AKSallocation. Create or update VMSS /subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j8084480-svc-aks1/providers/Microsoft.Compute/virtualMachineScaleSets/aks-u64d8dsv61-24652102-vmss failed. Allocation failed. VM(s) with the following constraints cannot be allocated, because the condition is too restrictive. Please remove some constraints and try again. Constraints applied are:\\\\n - Availability Zone\\\\n - Differencing (Ephemeral) Disks\\\\n - Networking Constraints (such as Accelerated Networking or IPv6)\\\\n - VM Size\\\\n\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"OverconstrainedZonalAllocationRequest\\\",\\r\\n \\\"target\\\": \\\"0\\\",\\r\\n \\\"message\\\": \\\"Allocation failed. VM(s) with the following constraints cannot be allocated, because the condition is too restrictive. Please remove some constraints and try again. Constraints applied are:\\\\n - Availability Zone\\\\n - Differencing (Ephemeral) Disks\\\\n - Networking Constraints (such as Accelerated Networking or IPv6)\\\\n - VM Size\\\\n\\\"\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n}\"\n }\n ]\n }\n}\n--------------------------------------------------------------------------------\n"time=2026-08-24T13:31:41.774Z level=INFO msg="Running step." serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=cluster stamp=2 description="Step cluster\n Kind: ARM\n Template: templates/mgmt-cluster.bicep\n Parameters: configurations/mgmt-cluster.tmpl.bicepparam"
time=2026-08-24T13:31:42.863Z level=DEBUG+2 msg="Starting ARM deployment" serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=cluster stamp=2
time=2026-08-24T13:31:44.020Z level=DEBUG+3 msg="Deployment started" serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=cluster stamp=2 deployment=cf8b4dcb00b286eb41bc03071739c50d179926b7125e9927b43e768ae20b10a5 portal=https://ms.portal.azure.com/#view/Microsoft_Azure_Resources/DeploymentDetails.MenuView/~/overview/id/%2Fsubscriptions%2FXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX%2FresourceGroups%2Fhcp-underlay-ci01-j5557248-mgmt-2%2Fproviders%2FMicrosoft.Resources%2Fdeployments%2Fcf8b4dcb00b286eb41bc03071739c50d179926b7125e9927b43e768ae20b10a5
time=2026-08-24T13:39:15.842Z level=ERROR msg="Step errored." serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=cluster stamp=2 err="stamp 2: failed to run ARM step: failed to poll deployment: failed to wait for deployment completion: GET https://management.azure.com/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j5557248-mgmt-2/providers/Microsoft.Resources/deployments/cf8b4dcb00b286eb41bc03071739c50d179926b7125e9927b43e768ae20b10a5/operationStatuses/08584140285820210806\n--------------------------------------------------------------------------------\nRESPONSE 200: 200 OK\nERROR CODE: DeploymentFailed\n--------------------------------------------------------------------------------\n{\n \"status\": \"Failed\",\n \"error\": {\n \"code\": \"DeploymentFailed\",\n \"message\": \"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\",\n \"details\": [\n {\n \"code\": \"Conflict\",\n \"message\": \"{\\r\\n \\\"status\\\": \\\"Failed\\\",\\r\\n \\\"error\\\": {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"DeploymentFailed\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j5557248-mgmt-2/providers/Microsoft.Resources/deployments/cluster-ce5kkcjcnxbu2\\\",\\r\\n \\\"message\\\": \\\"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j5557248-mgmt-2/providers/Microsoft.Resources/deployments/infra-agent-pools\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"DeploymentFailed\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j5557248-mgmt-2/providers/Microsoft.Resources/deployments/infra-agent-pools\\\",\\r\\n \\\"message\\\": \\\"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j5557248-mgmt-2/providers/Microsoft.ContainerService/managedClusters/ci01-j5557248-mgmt-2/agentPools/infra1\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"OverconstrainedZonalAllocationRequest\\\",\\r\\n \\\"message\\\": \\\"Allocation failed on resource '/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j5557248-mgmt-2-aks1/providers/Microsoft.Compute/virtualMachineScaleSets/aks-infra1-20013346-vmss'. Please try alternative sizes 'Standard_D8s_v6' (zones 3, 2), 'Standard_E8-4s_v6' (zones 3, 2), 'Standard_E8-4ds_v6' (zones 2, 3) instead for higher likelihood of success. Read more about improving likelihood of operation success at https://aka.ms/AKSallocation. Create or update VMSS /subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j5557248-mgmt-2-aks1/providers/Microsoft.Compute/virtualMachineScaleSets/aks-infra1-20013346-vmss failed. Allocation failed. VM(s) with the following constraints cannot be allocated, because the condition is too restrictive. Please remove some constraints and try again. Constraints applied are:\\\\n - Availability Zone\\\\n - Differencing (Ephemeral) Disks\\\\n - Networking Constraints (such as Accelerated Networking or IPv6)\\\\n - VM Size\\\\n\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"OverconstrainedZonalAllocationRequest\\\",\\r\\n \\\"target\\\": \\\"0\\\",\\r\\n \\\"message\\\": \\\"Allocation failed. VM(s) with the following constraints cannot be allocated, because the condition is too restrictive. Please remove some constraints and try again. Constraints applied are:\\\\n - Availability Zone\\\\n - Differencing (Ephemeral) Disks\\\\n - Networking Constraints (such as Accelerated Networking or IPv6)\\\\n - VM Size\\\\n\\\"\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n}\"\n }\n ]\n }\n}\n--------------------------------------------------------------------------------\n"time=2026-08-24T13:26:28.921Z level=INFO msg="Running step." serviceGroup=Microsoft.Azure.ARO.HCP.Service.Infra resourceGroup=service step=cluster description="Step cluster\n Kind: ARM\n Template: templates/svc-cluster.bicep\n Parameters: configurations/svc-cluster.tmpl.bicepparam"
time=2026-08-24T13:26:30.957Z level=DEBUG+2 msg="Starting ARM deployment" serviceGroup=Microsoft.Azure.ARO.HCP.Service.Infra resourceGroup=service step=cluster
time=2026-08-24T13:26:32.602Z level=DEBUG+3 msg="Deployment started" serviceGroup=Microsoft.Azure.ARO.HCP.Service.Infra resourceGroup=service step=cluster deployment=9508f110ddf9b49e79a1c674ffa6223d05182dbd61375e2ba689c04c8113e848 portal=https://ms.portal.azure.com/#view/Microsoft_Azure_Resources/DeploymentDetails.MenuView/~/overview/id/%2Fsubscriptions%2FXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX%2FresourceGroups%2Fhcp-underlay-ci01-j2132480-svc%2Fproviders%2FMicrosoft.Resources%2Fdeployments%2F9508f110ddf9b49e79a1c674ffa6223d05182dbd61375e2ba689c04c8113e848
time=2026-08-24T13:34:04.585Z level=ERROR msg="Step errored." serviceGroup=Microsoft.Azure.ARO.HCP.Service.Infra resourceGroup=service step=cluster err="failed to run ARM step: failed to poll deployment: failed to wait for deployment completion: GET https://management.azure.com/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j2132480-svc/providers/Microsoft.Resources/deployments/9508f110ddf9b49e79a1c674ffa6223d05182dbd61375e2ba689c04c8113e848/operationStatuses/08584140288938416315\n--------------------------------------------------------------------------------\nRESPONSE 200: 200 OK\nERROR CODE: DeploymentFailed\n--------------------------------------------------------------------------------\n{\n \"status\": \"Failed\",\n \"error\": {\n \"code\": \"DeploymentFailed\",\n \"message\": \"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\",\n \"details\": [\n {\n \"code\": \"Conflict\",\n \"message\": \"{\\r\\n \\\"status\\\": \\\"Failed\\\",\\r\\n \\\"error\\\": {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"DeploymentFailed\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j2132480-svc/providers/Microsoft.Resources/deployments/cluster-d54rkmqgu5xlq\\\",\\r\\n \\\"message\\\": \\\"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j2132480-svc/providers/Microsoft.Resources/deployments/user-agent-pools\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"DeploymentFailed\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j2132480-svc/providers/Microsoft.Resources/deployments/user-agent-pools\\\",\\r\\n \\\"message\\\": \\\"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j2132480-svc/providers/Microsoft.ContainerService/managedClusters/ci01-j2132480-svc/agentPools/u64d8dsv61\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"OverconstrainedZonalAllocationRequest\\\",\\r\\n \\\"message\\\": \\\"Allocation failed on resource '/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j2132480-svc-aks1/providers/Microsoft.Compute/virtualMachineScaleSets/aks-u64d8dsv61-42034125-vmss'. Please try alternative sizes 'Standard_D8s_v6' (zones 3, 2), 'Standard_E8-4s_v6' (zones 3, 2), 'Standard_E8-4ds_v6' (zones 2, 3) instead for higher likelihood of success. Read more about improving likelihood of operation success at https://aka.ms/AKSallocation. Create or update VMSS /subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j2132480-svc-aks1/providers/Microsoft.Compute/virtualMachineScaleSets/aks-u64d8dsv61-42034125-vmss failed. Allocation failed. VM(s) with the following constraints cannot be allocated, because the condition is too restrictive. Please remove some constraints and try again. Constraints applied are:\\\\n - Availability Zone\\\\n - Differencing (Ephemeral) Disks\\\\n - Networking Constraints (such as Accelerated Networking or IPv6)\\\\n - VM Size\\\\n\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"OverconstrainedZonalAllocationRequest\\\",\\r\\n \\\"target\\\": \\\"0\\\",\\r\\n \\\"message\\\": \\\"Allocation failed. VM(s) with the following constraints cannot be allocated, because the condition is too restrictive. Please remove some constraints and try again. Constraints applied are:\\\\n - Availability Zone\\\\n - Differencing (Ephemeral) Disks\\\\n - Networking Constraints (such as Accelerated Networking or IPv6)\\\\n - VM Size\\\\n\\\"\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n}\"\n }\n ]\n }\n}\n--------------------------------------------------------------------------------\n"Contributing tests (3)
Affected runs (28) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; host...full failure pattern: ERROR CODE: InternalServerError; detail message [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; hosted cluster degraded: UnavailableReplicas: [router]; provider Microsoft.RedHatOpenShift Signal: IndeterminateSignal: Indeterminate — no prior history; active 1 day(s) — New failure pattern — no prior history | e2e | 4 | 3.88%4 of 103 job runs affected | IndeterminateSignal: Indeterminate — no prior history; active 1 day(s) — New failure pattern — no prior history | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 10 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (3)fail [github.com/Azure/ARO-HCP/test/e2e/complete_cluster_create_multiversion.go:123]: HCP cluster rg-candidate-4-20-22lvs576ptrb/cluster-candidate-4-20-qdcf69 should provision
Unexpected error:
<*fmt.wrapError | 0xc001190ba0>:
failed to create HCP cluster cluster-candidate-4-20-qdcf69: failed waiting for cluster="cluster-candidate-4-20-qdcf69" in resourcegroup="rg-candidate-4-20-22lvs576ptrb" to finish creating: GET https://rp.j9409536.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/dc1dd563-6c7e-4d20-8fce-99e12ec22770
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/dc1dd563-6c7e-4d20-8fce-99e12ec22770",
"name": "dc1dd563-6c7e-4d20-8fce-99e12ec22770",
"status": "Failed",
"startTime": "2026-08-24T21:51:04.562343833Z",
"endTime": "2026-08-24T22:10:06.652345811Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; hosted cluster degraded: UnavailableReplicas: router deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/complete_cluster_create_multiversion.go:123]: HCP cluster rg-candidate-4-20-22lvs576ptrb/cluster-candidate-4-20-qdcf69 should provision
Unexpected error:
<*fmt.wrapError | 0xc001190ba0>:
failed to create HCP cluster cluster-candidate-4-20-qdcf69: failed waiting for cluster="cluster-candidate-4-20-qdcf69" in resourcegroup="rg-candidate-4-20-22lvs576ptrb" to finish creating: GET https://rp.j9409536.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/dc1dd563-6c7e-4d20-8fce-99e12ec22770
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/dc1dd563-6c7e-4d20-8fce-99e12ec22770",
"name": "dc1dd563-6c7e-4d20-8fce-99e12ec22770",
"status": "Failed",
"startTime": "2026-08-24T21:51:04.562343833Z",
"endTime": "2026-08-24T22:10:06.652345811Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; hosted cluster degraded: UnavailableReplicas: router deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_private_ingress.go:103]: failed to create HCP cluster "private-ingress" with private ingress
Unexpected error:
<*fmt.wrapError | 0xc000e33220>:
failed to create HCP cluster private-ingress: failed waiting for cluster="private-ingress" in resourcegroup="private-ingress-ndmwcdqgk685" to finish creating: GET https://rp.j9409536.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/a097060e-673b-40de-a82a-aa623235d88d
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/a097060e-673b-40de-a82a-aa623235d88d",
"name": "a097060e-673b-40de-a82a-aa623235d88d",
"status": "Failed",
"startTime": "2026-08-24T21:51:06.650394368Z",
"endTime": "2026-08-24T22:10:06.638990375Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; hosted cluster degraded: UnavailableReplicas: router deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_private_ingress.go:103]: failed to create HCP cluster "private-ingress" with private ingress
Unexpected error:
<*fmt.wrapError | 0xc000e33220>:
failed to create HCP cluster private-ingress: failed waiting for cluster="private-ingress" in resourcegroup="private-ingress-ndmwcdqgk685" to finish creating: GET https://rp.j9409536.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/a097060e-673b-40de-a82a-aa623235d88d
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/a097060e-673b-40de-a82a-aa623235d88d",
"name": "a097060e-673b-40de-a82a-aa623235d88d",
"status": "Failed",
"startTime": "2026-08-24T21:51:06.650394368Z",
"endTime": "2026-08-24T22:10:06.638990375Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; hosted cluster degraded: UnavailableReplicas: router deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/nodepool_version_upgrade.go:148]: failed to create HCP cluster np-version-upgrade-cluster-vkwj8z with version 4.21.30
Unexpected error:
<*fmt.wrapError | 0xc000bca9c0>:
failed to create HCP cluster np-version-upgrade-cluster-vkwj8z: failed waiting for cluster="np-version-upgrade-cluster-vkwj8z" in resourcegroup="rg-np-version-upgrade-vkwj8z-pvn2kjtfhwm4" to finish creating: GET https://rp.j6887040.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/9ca593c6-172c-442e-bdb4-2b96ae721c04
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/9ca593c6-172c-442e-bdb4-2b96ae721c04",
"name": "9ca593c6-172c-442e-bdb4-2b96ae721c04",
"status": "Failed",
"startTime": "2026-08-24T21:29:02.41159338Z",
"endTime": "2026-08-24T21:48:06.346460728Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; hosted cluster degraded: UnavailableReplicas: router deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/nodepool_version_upgrade.go:148]: failed to create HCP cluster np-version-upgrade-cluster-vkwj8z with version 4.21.30
Unexpected error:
<*fmt.wrapError | 0xc000bca9c0>:
failed to create HCP cluster np-version-upgrade-cluster-vkwj8z: failed waiting for cluster="np-version-upgrade-cluster-vkwj8z" in resourcegroup="rg-np-version-upgrade-vkwj8z-pvn2kjtfhwm4" to finish creating: GET https://rp.j6887040.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/9ca593c6-172c-442e-bdb4-2b96ae721c04
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/9ca593c6-172c-442e-bdb4-2b96ae721c04",
"name": "9ca593c6-172c-442e-bdb4-2b96ae721c04",
"status": "Failed",
"startTime": "2026-08-24T21:29:02.41159338Z",
"endTime": "2026-08-24T21:48:06.346460728Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; hosted cluster degraded: UnavailableReplicas: router deployment has 1 unavailable replicas"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (8)
Affected runs (4)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [svc] BackendControllerRetryHotLoop firedfull failure pattern: alert [svc] BackendControllerRetryHotLoop fired Signal: IndeterminateSignal: Indeterminate — seen in 3 prior week(s); active 2 day(s) | alert | 3 | 2.91%3 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 3 prior week(s); active 2 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 41 · Aug 22: 0 · Aug 23: 0 · Aug 24: 3 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (3)alert fired 1 time(s) Firing 1: State: Resolved Started: 2026-08-25T00:10:04Z Ended: 2026-08-25T00:34:04Z Severity: Sev3 Labels: alertname="BackendControllerRetryHotLoop", cluster="ci01-j8026752-svc", component="backend", name="observemanagedresourcegroup", severity="warning" Description: Backend controller workqueue observemanagedresourcegroup has a retry ratio of > 50% sustained over 10 minutes, indicating most queue activity is failed retries rather than fresh work. alert fired 1 time(s) Firing 1: State: Resolved Started: 2026-08-24T22:22:12Z Ended: 2026-08-24T22:46:11Z Severity: Sev3 Labels: alertname="BackendControllerRetryHotLoop", cluster="ci01-j2642688-svc", component="backend", name="observemanagedresourcegroup", severity="warning" Description: Backend controller workqueue observemanagedresourcegroup has a retry ratio of > 50% sustained over 10 minutes, indicating most queue activity is failed retries rather than fresh work. alert fired 1 time(s) Firing 1: State: Resolved Started: 2026-08-24T19:54:46Z Ended: 2026-08-24T20:17:49Z Severity: Sev3 Labels: alertname="BackendControllerRetryHotLoop", cluster="ci01-j4068096-svc", component="backend", name="observemanagedresourcegroup", severity="warning" Description: Backend controller workqueue observemanagedresourcegroup has a retry ratio of > 50% sustained over 10 minutes, indicating most queue activity is failed retries rather than fresh work. Contributing tests (1)
Affected runs (3)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
DNS for route host <route-host> did not resolve: DNS for <route-host> did not resolve within <duration> (last error: lookup <host> on <dns-server>: dial udp <ip>: i/o timeout): co...full failure pattern: DNS for route host <route-host> did not resolve: DNS for <route-host> did not resolve within <duration> (last error: lookup <host> on <dns-server>: dial udp <ip>: i/o timeout): context deadline exceeded Signal: IndeterminateSignal: Indeterminate — seen in 2 prior week(s); active 1 day(s) | e2e | 2 | 1.94%2 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 2 prior week(s); active 1 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 2 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (2)fail [github.com/Azure/ARO-HCP/test/e2e/complete_cluster_create_multiversion.go:189]: failed to verify simple web app runs on cluster "cluster-candidate-5-0-fn4s2m"
Unexpected error:
<*fmt.wrapError | 0xc000dea2e0>:
DNS for route host agnhost-e2e-sample-app-jhj46.apps.aro.o4q2a9i7j5z2k6t.i9i4.j4428544.hcp.osadev.cloud did not resolve: DNS for agnhost-e2e-sample-app-jhj46.apps.aro.o4q2a9i7j5z2k6t.i9i4.j4428544.hcp.osadev.cloud did not resolve within 10m0s (last error: lookup agnhost-e2e-sample-app-jhj46.apps.aro.o4q2a9i7j5z2k6t.i9i4.j4428544.hcp.osadev.cloud on 172.30.0.10:53: dial udp 172.30.0.10:53: i/o timeout): context deadline exceeded
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/complete_cluster_create_multiversion.go:189]: failed to verify simple web app runs on cluster "cluster-candidate-5-0-fn4s2m"
Unexpected error:
<*fmt.wrapError | 0xc000dea2e0>:
DNS for route host agnhost-e2e-sample-app-jhj46.apps.aro.o4q2a9i7j5z2k6t.i9i4.j4428544.hcp.osadev.cloud did not resolve: DNS for agnhost-e2e-sample-app-jhj46.apps.aro.o4q2a9i7j5z2k6t.i9i4.j4428544.hcp.osadev.cloud did not resolve within 10m0s (last error: lookup agnhost-e2e-sample-app-jhj46.apps.aro.o4q2a9i7j5z2k6t.i9i4.j4428544.hcp.osadev.cloud on 172.30.0.10:53: dial udp 172.30.0.10:53: i/o timeout): context deadline exceeded
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_v20260901preview.go:131]: failed to verify simple web app runs on v20260901preview cluster "v20260901"
Unexpected error:
<*fmt.wrapError | 0xc0005da4a0>:
DNS for route host agnhost-e2e-sample-app-gr4x8.apps.aro.v20260901.3e97.j6642304.hcp.osadev.cloud did not resolve: DNS for agnhost-e2e-sample-app-gr4x8.apps.aro.v20260901.3e97.j6642304.hcp.osadev.cloud did not resolve within 10m0s (last error: lookup agnhost-e2e-sample-app-gr4x8.apps.aro.v20260901.3e97.j6642304.hcp.osadev.cloud on 172.30.0.10:53: dial udp 172.30.0.10:53: i/o timeout): context deadline exceeded
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_v20260901preview.go:131]: failed to verify simple web app runs on v20260901preview cluster "v20260901"
Unexpected error:
<*fmt.wrapError | 0xc0005da4a0>:
DNS for route host agnhost-e2e-sample-app-gr4x8.apps.aro.v20260901.3e97.j6642304.hcp.osadev.cloud did not resolve: DNS for agnhost-e2e-sample-app-gr4x8.apps.aro.v20260901.3e97.j6642304.hcp.osadev.cloud did not resolve within 10m0s (last error: lookup agnhost-e2e-sample-app-gr4x8.apps.aro.v20260901.3e97.j6642304.hcp.osadev.cloud on 172.30.0.10:53: dial udp 172.30.0.10:53: i/o timeout): context deadline exceeded
...
occurredContributing tests (2)
Affected runs (2)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [svc] FrontendPathMedianLatency firedfull failure pattern: alert [svc] FrontendPathMedianLatency fired Signal: IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 2 day(s) | alert | 2 | 1.94%2 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 2 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 1 · Aug 23: 0 · Aug 24: 2 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (2)alert fired 1 time(s)
Firing 1:
State: Resolved
Started: 2026-08-24T17:55:35Z
Ended: 2026-08-24T18:27:33Z
Severity: Sev3
Labels: alertname="FrontendPathMedianLatency", cluster="ci01-j6361088-svc", component="frontend", method="put", route="/subscriptions/{subscriptionid}/resourcegroups/{resourcegroupname}/providers/microsoft.redhatopenshift/hcpopenshiftclusters/{resourcename}", severity="warning"
Description: The median (p50) of frontend request latency for put /subscriptions/{subscriptionid}/resourcegroups/{resourcegroupname}/providers/microsoft.redhatopenshift/hcpopenshiftclusters/{resourcename} has exceeded 250ms over the past 30 minutes (cluster ci01-j6361088-svc).alert fired 1 time(s)
Firing 1:
State: Resolved
Started: 2026-08-24T17:21:19Z
Ended: 2026-08-24T17:40:17Z
Severity: Sev3
Labels: alertname="FrontendPathMedianLatency", cluster="ci01-j9342336-svc", component="frontend", method="put", route="/subscriptions/{subscriptionid}/resourcegroups/{resourcegroupname}/providers/microsoft.redhatopenshift/hcpopenshiftclusters/{resourcename}", severity="warning"
Description: The median (p50) of frontend request latency for put /subscriptions/{subscriptionid}/resourcegroups/{resourcegroupname}/providers/microsoft.redhatopenshift/hcpopenshiftclusters/{resourcename} has exceeded 250ms over the past 30 minutes (cluster ci01-j9342336-svc).Contributing tests (1)
Affected runs (2)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: KubeconfigWaitingForCreate: ...full failure pattern: ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: KubeconfigWaitingForCreate: Waiting for hosted control plane kubeconfig to be created; hosted cluster degraded: UnavailableReplicas: [capi-provider, router]; provider Microsoft.RedHatOpenShift Signal: IndeterminateSignal: Indeterminate — seen in 4 prior week(s); active 2 day(s) | e2e | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 4 prior week(s); active 2 day(s) | Show trend detailsAug 18: 0 · Aug 19: 2 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 2 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (2)fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_private_kv.go:98]: failed to create HCP cluster "private-kv-cluster" with private keyvault
Unexpected error:
<*fmt.wrapError | 0xc000e70e20>:
failed waiting for cluster="private-kv-cluster" in resourcegroup="private-keyvault-wmkndtkx988t" to finish creating: GET https://rp.j9565440.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/a9d4baef-549e-4561-b748-439453d113af
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/a9d4baef-549e-4561-b748-439453d113af",
"name": "a9d4baef-549e-4561-b748-439453d113af",
"status": "Failed",
"startTime": "2026-08-24T15:49:08.39871397Z",
"endTime": "2026-08-24T16:08:11.236941815Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: KubeconfigWaitingForCreate: Waiting for hosted control plane kubeconfig to be created; hosted cluster degraded: UnavailableReplicas: [capi-provider deployment has 1 unavailable replicas, router deployment has 1 unavailable replicas]"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_private_kv.go:98]: failed to create HCP cluster "private-kv-cluster" with private keyvault
Unexpected error:
<*fmt.wrapError | 0xc000e70e20>:
failed waiting for cluster="private-kv-cluster" in resourcegroup="private-keyvault-wmkndtkx988t" to finish creating: GET https://rp.j9565440.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/a9d4baef-549e-4561-b748-439453d113af
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/a9d4baef-549e-4561-b748-439453d113af",
"name": "a9d4baef-549e-4561-b748-439453d113af",
"status": "Failed",
"startTime": "2026-08-24T15:49:08.39871397Z",
"endTime": "2026-08-24T16:08:11.236941815Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: KubeconfigWaitingForCreate: Waiting for hosted control plane kubeconfig to be created; hosted cluster degraded: UnavailableReplicas: [capi-provider deployment has 1 unavailable replicas, router deployment has 1 unavailable replicas]"
}
}
--------------------------------------------------------------------------------
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/cluster_fips_mode.go:106]: failed to create HCP cluster "fips-enabled-cluster" with cryptoRestrictions set to FIPS
Unexpected error:
<*fmt.wrapError | 0xc000408ec0>:
failed waiting for cluster="fips-enabled-cluster" in resourcegroup="fips-enabled-bn9cdfvxcqll" to finish creating: GET https://rp.j9565440.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/75726762-be88-402f-a564-60c91125dc8b
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/75726762-be88-402f-a564-60c91125dc8b",
"name": "75726762-be88-402f-a564-60c91125dc8b",
"status": "Failed",
"startTime": "2026-08-24T15:47:28.10624954Z",
"endTime": "2026-08-24T16:06:31.149377863Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: KubeconfigWaitingForCreate: Waiting for hosted control plane kubeconfig to be created; hosted cluster degraded: UnavailableReplicas: [capi-provider deployment has 1 unavailable replicas, router deployment has 1 unavailable replicas]"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_fips_mode.go:106]: failed to create HCP cluster "fips-enabled-cluster" with cryptoRestrictions set to FIPS
Unexpected error:
<*fmt.wrapError | 0xc000408ec0>:
failed waiting for cluster="fips-enabled-cluster" in resourcegroup="fips-enabled-bn9cdfvxcqll" to finish creating: GET https://rp.j9565440.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/75726762-be88-402f-a564-60c91125dc8b
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/75726762-be88-402f-a564-60c91125dc8b",
"name": "75726762-be88-402f-a564-60c91125dc8b",
"status": "Failed",
"startTime": "2026-08-24T15:47:28.10624954Z",
"endTime": "2026-08-24T16:06:31.149377863Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: KubeconfigWaitingForCreate: Waiting for hosted control plane kubeconfig to be created; hosted cluster degraded: UnavailableReplicas: [capi-provider deployment has 1 unavailable replicas, router deployment has 1 unavailable replicas]"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (2)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable:...full failure pattern: ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; provider Microsoft.RedHatOpenShift Signal: IndeterminateSignal: Indeterminate — seen in 2 prior week(s); active 1 day(s) | e2e | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 2 prior week(s); active 1 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 2 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (2)fail [github.com/Azure/ARO-HCP/test/e2e/nodepool_version_upgrade.go:304]: failed to create HCP cluster np-noedge-7pzvlw
Unexpected error:
<*fmt.wrapError | 0xc000e28760>:
failed to create HCP cluster np-noedge-7pzvlw: failed waiting for cluster="np-noedge-7pzvlw" in resourcegroup="rg-np-noedge-7pzvlw-hmhh49x6v5k9" to finish creating: GET https://rp.j6887040.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/d4851c53-78b5-4c3c-8fb8-61c89ed0b4a3
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/d4851c53-78b5-4c3c-8fb8-61c89ed0b4a3",
"name": "d4851c53-78b5-4c3c-8fb8-61c89ed0b4a3",
"status": "Failed",
"startTime": "2026-08-24T21:29:01.361816726Z",
"endTime": "2026-08-24T21:48:06.132587418Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/nodepool_version_upgrade.go:304]: failed to create HCP cluster np-noedge-7pzvlw
Unexpected error:
<*fmt.wrapError | 0xc000e28760>:
failed to create HCP cluster np-noedge-7pzvlw: failed waiting for cluster="np-noedge-7pzvlw" in resourcegroup="rg-np-noedge-7pzvlw-hmhh49x6v5k9" to finish creating: GET https://rp.j6887040.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/d4851c53-78b5-4c3c-8fb8-61c89ed0b4a3
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/d4851c53-78b5-4c3c-8fb8-61c89ed0b4a3",
"name": "d4851c53-78b5-4c3c-8fb8-61c89ed0b4a3",
"status": "Failed",
"startTime": "2026-08-24T21:29:01.361816726Z",
"endTime": "2026-08-24T21:48:06.132587418Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted"
}
}
--------------------------------------------------------------------------------
...
occurredfail [github.com/Azure/ARO-HCP/test/e2e/external_auth_create.go:95]: failed to create HCP cluster for external auth test
Unexpected error:
<*fmt.wrapError | 0xc000e95be0>:
failed to create HCP cluster ea-cluster: failed waiting for cluster="ea-cluster" in resourcegroup="external-auth-cluster-7w4qb7bx58wj" to finish creating: GET https://rp.j6887040.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/7c4e3b7e-7834-4bf7-9d07-9139f539790f
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/7c4e3b7e-7834-4bf7-9d07-9139f539790f",
"name": "7c4e3b7e-7834-4bf7-9d07-9139f539790f",
"status": "Failed",
"startTime": "2026-08-24T21:28:50.851778461Z",
"endTime": "2026-08-24T21:47:56.22619205Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/external_auth_create.go:95]: failed to create HCP cluster for external auth test
Unexpected error:
<*fmt.wrapError | 0xc000e95be0>:
failed to create HCP cluster ea-cluster: failed waiting for cluster="ea-cluster" in resourcegroup="external-auth-cluster-7w4qb7bx58wj" to finish creating: GET https://rp.j6887040.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/7c4e3b7e-7834-4bf7-9d07-9139f539790f
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/7c4e3b7e-7834-4bf7-9d07-9139f539790f",
"name": "7c4e3b7e-7834-4bf7-9d07-9139f539790f",
"status": "Failed",
"startTime": "2026-08-24T21:28:50.851778461Z",
"endTime": "2026-08-24T21:47:56.22619205Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (2)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [svc] FleetControllerRetryHotLoop firedfull failure pattern: alert [svc] FleetControllerRetryHotLoop fired Signal: IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 3 day(s) | alert | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 3 day(s) | Show trend detailsAug 18: 3 · Aug 19: 1 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)alert fired 1 time(s) Firing 1: State: Resolved Started: 2026-08-24T17:38:07Z Ended: 2026-08-24T17:45:05Z Severity: Sev3 Labels: alertname="FleetControllerRetryHotLoop", cluster="ci01-j0629888-svc", component="fleet", name="clustersserviceregistrationcontroller", severity="warning" Description: Fleet controller workqueue clustersserviceregistrationcontroller has a retry ratio of > 50% sustained over 10 minutes, indicating most queue activity is failed retries rather than fresh work. Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [svc] KubeContainerOOMKilled firedfull failure pattern: alert [svc] KubeContainerOOMKilled fired Signal: IndeterminateSignal: Indeterminate — seen in 3 prior week(s); active 2 day(s) | alert | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 3 prior week(s); active 2 day(s) | Show trend detailsAug 18: 3 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)alert fired 4 time(s) Firing 1: State: Fired (not resolved) Started: 2026-08-24T16:04:45Z Severity: Sev3 Labels: alertname="KubeContainerOOMKilled", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="fluentbit", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="arobit", pod="arobit-forwarder-cxwt8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", reason="OOMKilled", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning", uid="e03984bb-ef95-4df0-aec9-95c39b0e3c41" Description: Container fluentbit in pod arobit/arobit-forwarder-cxwt8 on cluster ci01-j9565440-mgmt-2 has been OOMKilled. This indicates the container exceeded its memory limit and was terminated by the kernel. Firing 2: State: Fired (not resolved) Started: 2026-08-24T16:04:46Z Severity: Sev3 Labels: alertname="KubeContainerOOMKilled", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="fluentbit", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="arobit", pod="arobit-forwarder-jbtcb", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", reason="OOMKilled", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning", uid="1ec91183-7c07-473b-b1d1-9df325a61c81" Description: Container fluentbit in pod arobit/arobit-forwarder-jbtcb on cluster ci01-j9565440-mgmt-2 has been OOMKilled. This indicates the container exceeded its memory limit and was terminated by the kernel. Firing 3: State: Fired (not resolved) Started: 2026-08-24T16:04:46Z Severity: Sev3 Labels: alertname="KubeContainerOOMKilled", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="fluentbit", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="arobit", pod="arobit-forwarder-jbtcb", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", reason="OOMKilled", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning", uid="1ec91183-7c07-473b-b1d1-9df325a61c81" Description: Container fluentbit in pod arobit/arobit-forwarder-jbtcb on cluster ci01-j9565440-mgmt-2 has been OOMKilled. This indicates the container exceeded its memory limit and was terminated by the kernel. Firing 4: State: Fired (not resolved) Started: 2026-08-24T16:04:46Z Severity: Sev3 Labels: alertname="KubeContainerOOMKilled", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="fluentbit", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="arobit", pod="arobit-forwarder-cxwt8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", reason="OOMKilled", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning", uid="e03984bb-ef95-4df0-aec9-95c39b0e3c41" Description: Container fluentbit in pod arobit/arobit-forwarder-cxwt8 on cluster ci01-j9565440-mgmt-2 has been OOMKilled. This indicates the container exceeded its memory limit and was terminated by the kernel. Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Cluster provisioning failedfull failure pattern: Cluster provisioning failed Signal: IndeterminateSignal: Indeterminate — seen in 4 prior week(s); active 2 day(s) | e2e | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 4 prior week(s); active 2 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 1 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/e2e/admin_credential_lifecycle.go:158]: Cluster provisioning failed fail [github.com/Azure/ARO-HCP/test/e2e/admin_credential_lifecycle.go:158]: Cluster provisioning failed Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: DeploymentFailed; detail code AllocationFailed; detail message The VM allocation failed due to an internal error. Please retry later or try deploying to a different lo...full failure pattern: ERROR CODE: DeploymentFailed; detail code AllocationFailed; detail message The VM allocation failed due to an internal error. Please retry later or try deploying to a different location.; provider Microsoft.Compute Signal: IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 1 day(s) | provision | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 1 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)time=2026-08-24T19:10:07.135Z level=INFO msg="Running step." serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=cluster stamp=2 description="Step cluster\n Kind: ARM\n Template: templates/mgmt-cluster.bicep\n Parameters: configurations/mgmt-cluster.tmpl.bicepparam"
time=2026-08-24T19:10:08.078Z level=DEBUG+2 msg="Starting ARM deployment" serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=cluster stamp=2
time=2026-08-24T19:10:08.961Z level=DEBUG+3 msg="Deployment started" serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=cluster stamp=2 deployment=6b7c37d7d05e023c4460016a815a9ccce8683c2a1a50686369f95e5de1b45b05 portal=https://ms.portal.azure.com/#view/Microsoft_Azure_Resources/DeploymentDetails.MenuView/~/overview/id/%2Fsubscriptions%2FXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX%2FresourceGroups%2Fhcp-underlay-ci01-j6400512-mgmt-2%2Fproviders%2FMicrosoft.Resources%2Fdeployments%2F6b7c37d7d05e023c4460016a815a9ccce8683c2a1a50686369f95e5de1b45b05
time=2026-08-24T19:19:10.348Z level=ERROR msg="Step errored." serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=cluster stamp=2 err="stamp 2: failed to run ARM step: failed to poll deployment: failed to wait for deployment completion: GET https://management.azure.com/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j6400512-mgmt-2/providers/Microsoft.Resources/deployments/6b7c37d7d05e023c4460016a815a9ccce8683c2a1a50686369f95e5de1b45b05/operationStatuses/08584140082768889498\n--------------------------------------------------------------------------------\nRESPONSE 200: 200 OK\nERROR CODE: DeploymentFailed\n--------------------------------------------------------------------------------\n{\n \"status\": \"Failed\",\n \"error\": {\n \"code\": \"DeploymentFailed\",\n \"message\": \"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\",\n \"details\": [\n {\n \"code\": \"Conflict\",\n \"message\": \"{\\r\\n \\\"status\\\": \\\"Failed\\\",\\r\\n \\\"error\\\": {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"DeploymentFailed\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j6400512-mgmt-2/providers/Microsoft.Resources/deployments/cluster-3vizzf2ahtjlc\\\",\\r\\n \\\"message\\\": \\\"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j6400512-mgmt-2/providers/Microsoft.Resources/deployments/user-agent-pools\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"DeploymentFailed\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j6400512-mgmt-2/providers/Microsoft.Resources/deployments/user-agent-pools\\\",\\r\\n \\\"message\\\": \\\"At least one resource deployment operation failed. Please list deployment operations for details. Please see https://aka.ms/arm-deployment-operations for usage details.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"ResourceDeploymentFailure\\\",\\r\\n \\\"target\\\": \\\"/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j6400512-mgmt-2/providers/Microsoft.ContainerService/managedClusters/ci01-j6400512-mgmt-2/agentPools/userswft1\\\",\\r\\n \\\"message\\\": \\\"The resource write operation failed to complete successfully, because it reached terminal provisioning state 'Failed'.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"AllocationFailed\\\",\\r\\n \\\"message\\\": \\\"Allocation failed on resource '/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourceGroups/hcp-underlay-ci01-j6400512-mgmt-2-aks1/providers/Microsoft.Compute/virtualMachineScaleSets/aks-userswft1-38184404-vmss'. Please try alternative sizes 'Standard_E16ds_v6' (zones 2, 3, 1), 'Standard_E16ds_v6' (eastus, zones 1, 2, 3), 'Standard_E16ds_v6' (eastus2, zones 2, 1, 3) instead for higher likelihood of success. Read more about improving likelihood of operation success at https://aka.ms/AKSallocation.\\\",\\r\\n \\\"details\\\": [\\r\\n {\\r\\n \\\"code\\\": \\\"AllocationFailed\\\",\\r\\n \\\"target\\\": \\\"3\\\",\\r\\n \\\"message\\\": \\\"The VM allocation failed due to an internal error. Please retry later or try deploying to a different location.\\\"\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n ]\\r\\n }\\r\\n}\"\n }\n ]\n }\n}\n--------------------------------------------------------------------------------\n"Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message cluster deletion did not complete before the deadline; [clusterServiceDeletion] ClusterService cluster still exists; [descendantRes...full failure pattern: ERROR CODE: InternalServerError; detail message cluster deletion did not complete before the deadline; [clusterServiceDeletion] ClusterService cluster still exists; [descendantResources] remaining resources: nodePools, serviceProviderClusters; provider Microsoft.RedHatOpenShift Signal: IndeterminateSignal: Indeterminate — seen in 2 prior week(s); active 1 day(s) | e2e | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 2 prior week(s); active 1 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/util/framework/per_test_framework.go:293]: Unexpected error:
<*errors.joinError | 0xc000ebc2a0>:
failed to cleanup resource group: at least one hcp cluster failed to delete: failed waiting for hcpCluster="cluster-candidate-5-0-h4ldj6" in resourcegroup="rg-candidate-5-0-xjnldggtrt7m" to finish deleting: GET https://rp.j9012096.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/fa0102ea-068c-4907-92f2-69579c623cac
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/fa0102ea-068c-4907-92f2-69579c623cac",
"name": "fa0102ea-068c-4907-92f2-69579c623cac",
"status": "Failed",
"startTime": "2026-08-24T15:21:51.698922375Z",
"endTime": "2026-08-24T15:45:59.987785659Z",
"error": {
"code": "InternalServerError",
"message": "cluster deletion did not complete before the deadline; [clusterServiceDeletion] ClusterService cluster /api/aro_hcp/v1alpha1/clusters/2sd0tm5ibpilvjmdnr77hbqf9t255bon still exists (deletion dispatched at 2026-08-24T15:21:51Z); [descendantResources] remaining resources: 1 Microsoft.RedHatOpenShift/hcpOpenShiftClusters/nodePools, 1 microsoft.redhatopenshift/hcpopenshiftclusters/serviceProviderClusters"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/util/framework/per_test_framework.go:293]: Unexpected error:
<*errors.joinError | 0xc000ebc2a0>:
failed to cleanup resource group: at least one hcp cluster failed to delete: failed waiting for hcpCluster="cluster-candidate-5-0-h4ldj6" in resourcegroup="rg-candidate-5-0-xjnldggtrt7m" to finish deleting: GET https://rp.j9012096.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/fa0102ea-068c-4907-92f2-69579c623cac
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/fa0102ea-068c-4907-92f2-69579c623cac",
"name": "fa0102ea-068c-4907-92f2-69579c623cac",
"status": "Failed",
"startTime": "2026-08-24T15:21:51.698922375Z",
"endTime": "2026-08-24T15:45:59.987785659Z",
"error": {
"code": "InternalServerError",
"message": "cluster deletion did not complete before the deadline; [clusterServiceDeletion] ClusterService cluster /api/aro_hcp/v1alpha1/clusters/2sd0tm5ibpilvjmdnr77hbqf9t255bon still exists (deletion dispatched at 2026-08-24T15:21:51Z); [descendantResources] remaining resources: 1 Microsoft.RedHatOpenShift/hcpOpenShiftClusters/nodePools, 1 microsoft.redhatopenshift/hcpopenshiftclusters/serviceProviderClusters"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster has no installed version; hosted cluster degra...full failure pattern: ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster has no installed version; hosted cluster degraded: UnavailableReplicas: [hosted-cluster-config-operator, ignition-server-proxy, kube-controller-manager, router]; provider Microsoft.RedHatOpenShift Signal: IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 2 day(s) | e2e | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 2 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 1 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_hypershift_presubmit.go:82]: HCP cluster rg-hs-pre--wkl8qfr9l4cs/cluster-hs-pre-jtp7jx should provision
Unexpected error:
<*fmt.wrapError | 0xc0009c0960>:
failed to create HCP cluster cluster-hs-pre-jtp7jx: failed waiting for cluster="cluster-hs-pre-jtp7jx" in resourcegroup="rg-hs-pre--wkl8qfr9l4cs" to finish creating: GET https://rp.j1362816.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/df253c01-fd54-4465-9725-47d0942d268e
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/df253c01-fd54-4465-9725-47d0942d268e",
"name": "df253c01-fd54-4465-9725-47d0942d268e",
"status": "Failed",
"startTime": "2026-08-24T15:18:48.695053949Z",
"endTime": "2026-08-24T15:37:57.993163664Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster has no installed version; hosted cluster degraded: UnavailableReplicas: [hosted-cluster-config-operator deployment has 1 unavailable replicas, ignition-server-proxy deployment has 1 unavailable replicas, kube-controller-manager deployment has 1 unavailable replicas, router deployment has 1 unavailable replicas]"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_hypershift_presubmit.go:82]: HCP cluster rg-hs-pre--wkl8qfr9l4cs/cluster-hs-pre-jtp7jx should provision
Unexpected error:
<*fmt.wrapError | 0xc0009c0960>:
failed to create HCP cluster cluster-hs-pre-jtp7jx: failed waiting for cluster="cluster-hs-pre-jtp7jx" in resourcegroup="rg-hs-pre--wkl8qfr9l4cs" to finish creating: GET https://rp.j1362816.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/df253c01-fd54-4465-9725-47d0942d268e
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/df253c01-fd54-4465-9725-47d0942d268e",
"name": "df253c01-fd54-4465-9725-47d0942d268e",
"status": "Failed",
"startTime": "2026-08-24T15:18:48.695053949Z",
"endTime": "2026-08-24T15:37:57.993163664Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster has no installed version; hosted cluster degraded: UnavailableReplicas: [hosted-cluster-config-operator deployment has 1 unavailable replicas, ignition-server-proxy deployment has 1 unavailable replicas, kube-controller-manager deployment has 1 unavailable replicas, router deployment has 1 unavailable replicas]"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable; host...full failure pattern: ERROR CODE: InternalServerError; detail message [clusterServiceClusterStatus] <no_message>; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable; hosted cluster degraded: UnavailableReplicas: [catalog-operator, hosted-cluster-config-operator, olm-operator, openshift-apiserver, packageserver]; provider Microsoft.RedHatOpenShift Signal: IndeterminateSignal: Indeterminate — seen in 2 prior week(s); active 2 day(s) | e2e | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 2 prior week(s); active 2 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 1 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_feature_aggregation.go:184]: failed to create HCP cluster "agg-cluster" with aggregated settings
Unexpected error:
<*fmt.wrapError | 0xc000a808e0>:
failed waiting for cluster="agg-cluster" in resourcegroup="feature-aggregation-cgt6b2gspgmg" to finish creating: GET https://rp.j1362816.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/c64a33f9-1d34-4ca6-89af-ad50a67745b6
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/c64a33f9-1d34-4ca6-89af-ad50a67745b6",
"name": "c64a33f9-1d34-4ca6-89af-ad50a67745b6",
"status": "Failed",
"startTime": "2026-08-24T15:19:10.873312133Z",
"endTime": "2026-08-24T15:38:17.935941947Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable: Waiting for components to be available: cluster-version-operator, cluster-node-tuning-operator, redhat-marketplace-catalog, cluster-network-operator, redhat-operators-catalog, catalog-operator, openshift-route-controller-manager, ingress-operator, hosted-cluster-config-operator, packageserver, cluster-policy-controller, cluster-storage-operator, certified-operators-catalog, olm-operator, openshift-controller-manager, csi-snapshot-controller-operator, community-operators-catalog, dns-operator; hosted cluster degraded: UnavailableReplicas: [catalog-operator deployment has 1 unavailable replicas, hosted-cluster-config-operator deployment has 1 unavailable replicas, olm-operator deployment has 1 unavailable replicas, openshift-apiserver deployment has 1 unavailable replicas, packageserver deployment has 3 unavailable replicas]"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/cluster_create_feature_aggregation.go:184]: failed to create HCP cluster "agg-cluster" with aggregated settings
Unexpected error:
<*fmt.wrapError | 0xc000a808e0>:
failed waiting for cluster="agg-cluster" in resourcegroup="feature-aggregation-cgt6b2gspgmg" to finish creating: GET https://rp.j1362816.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/c64a33f9-1d34-4ca6-89af-ad50a67745b6
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/c64a33f9-1d34-4ca6-89af-ad50a67745b6",
"name": "c64a33f9-1d34-4ca6-89af-ad50a67745b6",
"status": "Failed",
"startTime": "2026-08-24T15:19:10.873312133Z",
"endTime": "2026-08-24T15:38:17.935941947Z",
"error": {
"code": "InternalServerError",
"message": "[clusterServiceClusterStatus] \u003cno_message\u003e; [hypershiftHostedCluster] hosted cluster is not available: ComponentsNotAvailable: Waiting for components to be available: cluster-version-operator, cluster-node-tuning-operator, redhat-marketplace-catalog, cluster-network-operator, redhat-operators-catalog, catalog-operator, openshift-route-controller-manager, ingress-operator, hosted-cluster-config-operator, packageserver, cluster-policy-controller, cluster-storage-operator, certified-operators-catalog, olm-operator, openshift-controller-manager, csi-snapshot-controller-operator, community-operators-catalog, dns-operator; hosted cluster degraded: UnavailableReplicas: [catalog-operator deployment has 1 unavailable replicas, hosted-cluster-config-operator deployment has 1 unavailable replicas, olm-operator deployment has 1 unavailable replicas, openshift-apiserver deployment has 1 unavailable replicas, packageserver deployment has 3 unavailable replicas]"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
ERROR CODE: InternalServerError; detail message [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; prov...full failure pattern: ERROR CODE: InternalServerError; detail message [hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted; provider Microsoft.RedHatOpenShift Signal: IndeterminateSignal: Indeterminate — seen in 3 prior week(s); active 2 day(s) | e2e | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 3 prior week(s); active 2 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 1 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/e2e/complete_cluster_create_multiversion.go:123]: HCP cluster rg-candidate-4-21-cbhfzqtvb9th/cluster-candidate-4-21-5bbfxw should provision
Unexpected error:
<*fmt.wrapError | 0xc000ea2680>:
failed to create HCP cluster cluster-candidate-4-21-5bbfxw: failed waiting for cluster="cluster-candidate-4-21-5bbfxw" in resourcegroup="rg-candidate-4-21-cbhfzqtvb9th" to finish creating: GET https://rp.j2463616.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/a86d82f7-93d6-43a5-99ee-59e2969fd318
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/a86d82f7-93d6-43a5-99ee-59e2969fd318",
"name": "a86d82f7-93d6-43a5-99ee-59e2969fd318",
"status": "Failed",
"startTime": "2026-08-24T17:28:32.844651935Z",
"endTime": "2026-08-24T17:47:34.560561952Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted"
}
}
--------------------------------------------------------------------------------
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/complete_cluster_create_multiversion.go:123]: HCP cluster rg-candidate-4-21-cbhfzqtvb9th/cluster-candidate-4-21-5bbfxw should provision
Unexpected error:
<*fmt.wrapError | 0xc000ea2680>:
failed to create HCP cluster cluster-candidate-4-21-5bbfxw: failed waiting for cluster="cluster-candidate-4-21-5bbfxw" in resourcegroup="rg-candidate-4-21-cbhfzqtvb9th" to finish creating: GET https://rp.j2463616.hcpsvc.osadev.cloud/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/a86d82f7-93d6-43a5-99ee-59e2969fd318
--------------------------------------------------------------------------------
RESPONSE 200: 200 OK
ERROR CODE: InternalServerError
--------------------------------------------------------------------------------
{
"id": "/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/providers/Microsoft.RedHatOpenShift/locations/canadacentral/hcpOperationStatuses/a86d82f7-93d6-43a5-99ee-59e2969fd318",
"name": "a86d82f7-93d6-43a5-99ee-59e2969fd318",
"status": "Failed",
"startTime": "2026-08-24T17:28:32.844651935Z",
"endTime": "2026-08-24T17:47:34.560561952Z",
"error": {
"code": "InternalServerError",
"message": "[hypershiftHostedCluster] hosted cluster is not available: KASLoadBalancerNotReachable: APIServer external route not admitted"
}
}
--------------------------------------------------------------------------------
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
UpdateHCPCluster20260630 timed out after <minutes> minutes; context deadline exceededfull failure pattern: UpdateHCPCluster20260630 timed out after <minutes> minutes; context deadline exceeded Signal: IndeterminateSignal: Indeterminate — seen in 4 prior week(s); active 1 day(s) | e2e | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 4 prior week(s); active 1 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/e2e/kms_key_rotation.go:172]: failed to update cluster with new KMS key
Unexpected error:
<*fmt.wrapErrors | 0xc00132c390>:
failed waiting for hcpCluster="kms-key-rotate-422" in resourcegroup="kms-key-rotate-cflktt657pch" to finish updating, caused by: timeout '30.000000' minutes exceeded during UpdateHCPCluster20260630 for cluster kms-key-rotate-422 in resource group kms-key-rotate-cflktt657pch, error: context deadline exceeded
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/kms_key_rotation.go:172]: failed to update cluster with new KMS key
Unexpected error:
<*fmt.wrapErrors | 0xc00132c390>:
failed waiting for hcpCluster="kms-key-rotate-422" in resourcegroup="kms-key-rotate-cflktt657pch" to finish updating, caused by: timeout '30.000000' minutes exceeded during UpdateHCPCluster20260630 for cluster kms-key-rotate-422 in resource group kms-key-rotate-cflktt657pch, error: context deadline exceeded
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
UpdateNodePoolAndWait timed out after <minutes> minutes; context deadline exceededfull failure pattern: UpdateNodePoolAndWait timed out after <minutes> minutes; context deadline exceeded Signal: IndeterminateSignal: Indeterminate — seen in 4 prior week(s); active 1 day(s) | e2e | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 4 prior week(s); active 1 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/e2e/nodepool_update_nodes.go:208]: failed to enable autoscaling on node pool np-one-node
Unexpected error:
<*fmt.wrapErrors | 0xc00396d0e0>:
failed waiting for nodepool="np-one-node" in cluster="np-update-nodes-hcp-cluster" resourcegroup="nodepool-update-nodes-4984p9m7f6lc" to finish updating, caused by: timeout '20.000000' minutes exceeded during UpdateNodePoolAndWait for nodepool np-one-node in cluster np-update-nodes-hcp-cluster in resource group nodepool-update-nodes-4984p9m7f6lc, error: context deadline exceeded
...
occurred
fail [github.com/Azure/ARO-HCP/test/e2e/nodepool_update_nodes.go:208]: failed to enable autoscaling on node pool np-one-node
Unexpected error:
<*fmt.wrapErrors | 0xc00396d0e0>:
failed waiting for nodepool="np-one-node" in cluster="np-update-nodes-hcp-cluster" resourcegroup="nodepool-update-nodes-4984p9m7f6lc" to finish updating, caused by: timeout '20.000000' minutes exceeded during UpdateNodePoolAndWait for nodepool np-one-node in cluster np-update-nodes-hcp-cluster in resource group nodepool-update-nodes-4984p9m7f6lc, error: context deadline exceeded
...
occurredContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
alert [svc] KubePodCrashLooping firedfull failure pattern: alert [svc] KubePodCrashLooping fired Signal: IndeterminateSignal: Indeterminate — seen in 3 prior week(s); active 1 day(s) | alert | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 3 prior week(s); active 1 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)alert fired 2 time(s) Firing 1: State: Resolved Started: 2026-08-24T16:19:07Z Ended: 2026-08-24T16:30:05Z Severity: Sev3 Labels: alertname="KubePodCrashLooping", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="fluentbit", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="arobit", pod="arobit-forwarder-cxwt8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-0", reason="CrashLoopBackOff", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning", uid="e03984bb-ef95-4df0-aec9-95c39b0e3c41" Description: Pod arobit/arobit-forwarder-cxwt8 (fluentbit) is in waiting state (reason: "CrashLoopBackOff"). Firing 2: State: Resolved Started: 2026-08-24T16:20:06Z Ended: 2026-08-24T16:33:06Z Severity: Sev3 Labels: alertname="KubePodCrashLooping", cluster="ci01-j9565440-mgmt-2", component="kubernetes-infrastructure", container="fluentbit", endpoint="http", environment="ci01", instance="10.128.64.73:8080", job="kube-state-metrics", microsoft.amwresourceid="/subscriptions/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX/resourcegroups/hcp-underlay-ci01-j9565440/providers/microsoft.monitor/accounts/services-j9565440", namespace="arobit", pod="arobit-forwarder-cxwt8", prometheus="prometheus/prometheus", prometheus_replica="prom-agent-prometheus-1", reason="CrashLoopBackOff", region="canadacentral", service="arohcp-monitor-kube-state-metrics", severity="warning", uid="e03984bb-ef95-4df0-aec9-95c39b0e3c41" Description: Pod arobit/arobit-forwarder-cxwt8 (fluentbit) is in waiting state (reason: "CrashLoopBackOff"). Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
failed to run ARM step: failed to poll deployment: failed to wait for deployment completion: context deadline exceededfull failure pattern: failed to run ARM step: failed to poll deployment: failed to wait for deployment completion: context deadline exceeded Signal: IndeterminateSignal: Indeterminate — seen in 2 prior week(s); active 1 day(s) | provision | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 2 prior week(s); active 1 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)time=2026-08-24T09:44:43.919Z level=INFO msg="Running step." serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=cluster stamp=1 description="Step cluster\n Kind: ARM\n Template: templates/mgmt-cluster.bicep\n Parameters: configurations/mgmt-cluster.tmpl.bicepparam" time=2026-08-24T09:44:44.928Z level=DEBUG+2 msg="Starting ARM deployment" serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=cluster stamp=1 time=2026-08-24T09:44:46.104Z level=DEBUG+3 msg="Deployment started" serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=cluster stamp=1 deployment=f40d7b7ab0b2d9646dc8b78fbaa252859442c0d4d4549c8765d66346f06ac123 portal=https://ms.portal.azure.com/#view/Microsoft_Azure_Resources/DeploymentDetails.MenuView/~/overview/id/%2Fsubscriptions%2FXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX%2FresourceGroups%2Fhcp-underlay-ci01-j4585344-mgmt-1%2Fproviders%2FMicrosoft.Resources%2Fdeployments%2Ff40d7b7ab0b2d9646dc8b78fbaa252859442c0d4d4549c8765d66346f06ac123 time=2026-08-24T10:14:43.424Z level=ERROR msg="Step errored." serviceGroup=Microsoft.Azure.ARO.HCP.Management.Infra resourceGroup=management step=cluster stamp=1 err="stamp 1: failed to run ARM step: failed to poll deployment: failed to wait for deployment completion: context deadline exceeded" Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
failed to select node pool <nodepool> version for <ocp-channel> channel: Get "<url>": dial tcp: lookup apifull failure pattern: failed to select node pool <nodepool> version for <ocp-channel> channel: Get "<url>": dial tcp: lookup api Signal: IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 2 day(s) | e2e | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 2 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 1 · Aug 21: 0 · Aug 22: 0 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/util/framework/deployment_params.go:207]: failed to select node pool install version for candidate-4.20 channel: Get "https://api.openshift.com/api/upgrades_info/graph?arch=multi&channel=candidate-4.20": dial tcp: lookup api.openshift.com on 172.30.0.10:53: read udp 10.128.208.156:42093->172.30.0.10:53: i/o timeout fail [github.com/Azure/ARO-HCP/test/util/framework/deployment_params.go:207]: failed to select node pool install version for candidate-4.20 channel: Get "https://api.openshift.com/api/upgrades_info/graph?arch=multi&channel=candidate-4.20": dial tcp: lookup api.openshift.com on 172.30.0.10:53: read udp 10.128.208.156:42093->172.30.0.10:53: i/o timeout Contributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
stamp 2 scheduling must have requested resourcesfull failure pattern: stamp 2 scheduling must have requested resources Signal: IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 2 day(s) | e2e | 1 | 0.97%1 of 103 job runs affected | IndeterminateSignal: Indeterminate — seen in 1 prior week(s); active 2 day(s) | Show trend detailsAug 18: 0 · Aug 19: 0 · Aug 20: 0 · Aug 21: 0 · Aug 22: 1 · Aug 23: 0 · Aug 24: 1 | none | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Full failure examples (1)fail [github.com/Azure/ARO-HCP/test/e2e/fleet_registration.go:162]: Timed out after 900.001s.
scheduling data was not populated in time
The function passed to Eventually failed at /opt/app-root/src/github.com/Azure/ARO-HCP/test/e2e/fleet_registration.go:155 with:
stamp 2 scheduling must have requested resources
Expected
<v1.ResourceList | len:0>: ...
not to be empty
fail [github.com/Azure/ARO-HCP/test/e2e/fleet_registration.go:162]: Timed out after 900.001s.
scheduling data was not populated in time
The function passed to Eventually failed at /opt/app-root/src/github.com/Azure/ARO-HCP/test/e2e/fleet_registration.go:155 with:
stamp 2 scheduling must have requested resources
Expected
<v1.ResourceList | len:0>: ...
not to be emptyContributing tests (1)
Affected runs (1)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||