Teams in the financial services industry face a unique set of drivers for adopting observability tools.
- Regulatory compliance from DORA to MAS TRM gives strict guidelines for cyber incidents and data retention
- The incidents themselves – which observability tools hold the promise of reducing – are among the most costly for financial services compared to virtually any other vertical
- Gen AI, now widely adopted, creates a new meta layer of observation: who watches the watchers, we must ask.
In this piece, we look at 7 popular observability platforms: their capabilities, compliance posture, pricing, regulation fit, and what they are good for.
We work with FinTech and FinServ teams on observability, incident management and CloudOps. If you would rather talk it through than read to the end, you can contact us here.
Datadog
Capabilities
Datadog is the broadest single-vendor observability platform on this list, covering logs, metrics, traces, real user monitoring, synthetics and profiling with Cloud SIEM layered on the same data pipeline.
This enables a unified view from threat to infra which is useful for coordinating between security and engineering teams.
- Deployment: SaaS, primarily. Datadog’s CloudPrem adds the option to index and search logs in your own infrastructure, but only logs (metrics, traces and the control plane stay on Datadog) and this is an emerging offering. If your internal policy covers all telemetry, CloudPrem does not close the gap
- Retention: Of Datadog’s tiers, Standard indexes offer 3 to 60 days. Flex extends to 15 months, Flex Frozen to seven years without rehydration (still in private preview). You can also archive to your own object storage and rehydrate on demand
Compliance posture
- PCI DSS
- SOC 2 Type 2
- ISO 27001
- IRAP
- FedRAMP High (May 2026; covers Datadog for Government, not the commercial platform)
Pricing model
Modular and usage-based. You pay separately for ingestion (per GB) and indexing (per event), plus additional charges per product – APM, infrastructure monitoring, synthetics and so on each carry their own line item.
Longer retention and cheaper storage tiers are available, but querying against them costs extra. At the log volumes financial services environments generate (particularly if you are retaining transaction and audit data to meet PCI or DORA requirements), costs escalate quickly. The pricing model rewards careful data management.
Regulatory fit
- DORA: No hard blockers
- MAS TRM: Should work well. But worth noting the one-hour window on severe incident reporting means routing critical logs to Flex to save on costs (and losing out on real-time detection) may be very unwise
- FCA/PRA: Well suited to tolerance monitoring and scenario evidence. Under the new third-party reporting rules applying from March 2027, Datadog itself will sit on your register of material third-party arrangements
- APRA CPS 230: No blockers. Works for continuous monitoring and tolerance evidence
Best for
Financial institutions with budget and a platform team who will actively manage data tiering. A mid-to-large PSP or digital bank fits here. Not viable if all telemetry, not just logs, must stay inside your own infrastructure due to internal policy.
Dynatrace
Capabilities
Dynatrace’s distinguishing feature is its topology model. OneAgent maps dependencies automatically across the technology stack, and the Davis AI engine uses that model for causal root cause analysis rather than mere correlation. If you are running a payment platform with dozens of upstream and downstream integrations, that outsourced mental work can be crucial in a pinch.
- Deployment: SaaS or Managed (self-hosted). Both remain available and supported, but they are diverging. Grail, the platform’s long-term retention and query engine, is SaaS-only with no plans to bring to Managed users. The same applies to AppEngine, AutomationEngine and EdgeConnect. On SaaS, Grail provides 15 months of metric retention at one-minute granularity as standard, extensible to ten years. On Managed, retention is whatever storage you provision yourself
- Retention: That SaaS/Managed split can complicate things. If you opt for Managed to satisfy internal policies, you lose out on ease of use and the more powerful features of the platform
Compliance posture
- PCI DSS
- SOC 2 Type 2
- ISO 27001
- IRAP
- FedRAMP Moderate (announced intent to pursue FedRAMP High in July 2026; authorisation not yet achieved)
Pricing model
Consumption-based through the Dynatrace Platform Subscription, metered across monitoring, log management, Grail retention and Davis queries. Enterprise costs. Extended Grail retention is a purchasable line item on SaaS and simply unavailable on Managed.
Regulatory fit
- DORA: Strong on SaaS, where Davis and Grail between them cover both detection and evidence well. On Managed, detection is still strong but you are building the evidence layer yourself
- MAS TRM: Davis’s automated root cause analysis is well suited to fast incident identification. Same SaaS/Managed caveat applies to retention
- FCA/PRA: The topology model makes mapping important business services and their dependencies more straightforward than on most observability platforms. No blockers
- APRA CPS 230: No blockers. Dynatrace holds IRAP at the Protected level on both AWS and Azure, which simplifies procurement for Australian financial institutions
Best for
Large, complex FS estates: a bank or insurer with a deep microservices architecture and many third-party integrations where the root cause of an outage is rarely obvious. If in-region SaaS is acceptable, this is one of the strongest options.
If residency requirements push you towards Managed, make sure you understand what you are giving up. Again, Grail is not headed for Managed at the time of writing.
Elastic Observability
Capabilities
Elastic’s strength is flexibility; see what the marketing team did there? It ingests logs, metrics, traces and profiling; it is OpenTelemetry-native, and it runs the same query language and analytics across observability and security workloads, like Datadog.
- Deployment: The widest range on this list. Self-managed on your own infrastructure, Elastic Cloud Hosted across 50+ regions on AWS, Azure and Google Cloud, or Elastic Cloud Serverless. Self-managed means all telemetry stays inside your perimeter, not just logs. This makes it something of a standout offer
- Retention: In your hands. Hot, warm, cold and frozen tiers back onto object storage you control, and logsdb index mode reduces storage costs substantially for log-heavy workloads. There is no vendor-imposed ceiling on how long you keep data
Compliance posture
- PCI DSS
- SOC 2 Type 2
- ISO 27001
- IRAP
- FedRAMP High (Cloud Hosted, March 2026 – US government only)
Pricing model
Depends on deployment. Cloud hosted is resource-based, Serverless is consumption-based, and self-managed is licence plus your own infrastructure.
Regulatory fit
- DORA: Self-managed simplifies the third-party ICT provider dimension, since your vendor relationship reduces to a licence rather than a data processor. Full control over data location and retention
- MAS TRM: No blockers. Self-managed or in-region cloud-hosted both work.
- FCA/PRA: Works well for tolerance monitoring and evidence. Same third-party simplification as DORA if self-managed
- APRA CPS 230: The strongest fit on this list for Australian financial services specifically, because it combines IRAP with full self-managed deployment, removing the third-party data processor relationship
Best for
An FCA-regulated FinTech or APRA-regulated lender with strict residency requirements and a hands-on, well-staffed engineering team. Also a strong pick where security and observability should share a data layer.
Splunk (Cisco)
Capabilities
Splunk has the deepest log analytics heritage on this list, and it remains the strongest observability platform for dealing with messy data due to its history: it was built to and continues to be optimised for querying unstructured data.
- Deployment: Both SaaS (Splunk Cloud) and fully self-hosted (Splunk Enterprise). The self-hosted option is genuinely full-featured, not a limited subset, which is vital for some use cases
- Retention: On Splunk Enterprise, retention is bounded only by the storage you provision. On Splunk Cloud, standard retention tiers apply with longer retention available by arrangement. SmartStore separates compute from storage, which helps manage cost at the data volumes FS environments generate
Compliance posture
- PCI DSS
- SOC 2 Type 2
- ISO 27001
- IRAP
- FedRAMP High (Splunk Cloud)
Pricing model
Two models. Ingest-based pricing charges per GB per day. Straightforward but punishing if your data volumes are high relative to how much you query. Workload pricing, which Cisco now steers new customers towards, charges for compute consumed by searches and analytics rather than raw ingest.
It suits environments that ingest heavily but query selectively, which describes many FS compliance logging patterns. Splunk Enterprise Security, ITSI and SOAR each carry separate licences, and ES in particular can roughly double the effective cost of a deployment.
Regulatory fit
- DORA: No blockers on either deployment. Self-hosted Splunk Enterprise gives you the same third-party simplification as self-managed Elastic
- MAS TRM: Works well, particularly if you already use Splunk for security.
- FCA/PRA: Splunk is already embedded in many UK financial institutions for security analytics, so extending it to observability keeps things clean
- APRA CPS 230: No blockers. Splunk holds IRAP at the Protected level, with 20 IRAP-assessed offerings including Splunk Observability Cloud
Best for
Larger, older financial institutions that need observability in systems that output unstructured data. Fully featured self-hosting also points toward a larger, older and highly regulated ideal user. Less a draw for the nimble, upstart FinTech.
Grafana Cloud (LGTM stack)
Capabilities
Grafana Cloud is the managed version of the open-source LGTM stack, Loki for logs, Mimir for metrics, Tempo for traces, and Grafana for the UI. The appeal is openness: it is built on OpenTelemetry and Prometheus from the ground up, so there is less vendor lock-in than with any proprietary observability platform on this list. If you leave, your instrumentation and your query language (PromQL, LogQL, TraceQL) go with you.
- Deployment: SaaS on Grafana Cloud, or fully self-hosted using the open-source components. The self-hosted path gives you complete control over data location and retention, but you are running four separate backends and their operational overhead is substantial. Grafana Labs themselves have said a dedicated SRE team is realistic for production self-hosted deployments. So, not one to rush into (see this piece for a dive into deployment options)
- Retention: On Grafana Cloud, retention is configurable within the tiers your plan allows. Self-hosted, there is no ceiling
Compliance posture
- PCI DSS
- SOC 2 Type 2
- ISO 27001
- FedRAMP High (Federal Cloud)
Pricing model
Usage-based across separate meters for metrics (per active series), logs (per GB processed, written and retained) and traces (per GB). A free tier covers small workloads, and paid plans start at $19/month plus usage.
The metering is transparent but additive. Each backend bills independently, so what looks like one observability platform on the dashboard is several line items on the invoice. Self-hosted replaces all of that with infrastructure and engineering cost.
Regulatory fit
- DORA: Self-hosted LGTM gives full control over data location and retention with minimal vendor dependency. On Grafana Cloud, no hard blockers but you are introducing a data processor
- MAS TRM: No blockers on either deployment. The platform is less automated than Datadog or Dynatrace in its observability capabilities around detection, so you are more reliant on well-configured alerting rules.
- FCA/PRA: Works well, particularly self-hosted. Same third-party considerations as the other SaaS options if you use Grafana Cloud
- APRA CPS 230: No hard blockers, though Grafana Cloud does not natively hold IRAP
Best for
Engineering-led FS teams – a FinTech with strong platform engineering capability that wants to avoid vendor lock-in and is willing to invest the operational time to get there. Also a good fit if your tech stack already includes Prometheus and Grafana dashboards.
Sumo Logic
Capabilities
Sumo Logic sits between observability and SIEM, which is both its strength and its weakness. It is strong on log analytics and cloud-native security (Cloud SIEM, SOAR, UEBA) as this was its founding purpose. However, its observability features were added later, and are less mature than others.
- Deployment: SaaS only, but with a wider range of regional options than most SaaS-only observability tools. Commercial deployments are available in eight regions, including Sydney, Tokyo and Frankfurt, to name a few
- Retention: Configurable across hot, cold and frozen tiers. Tamper-evident audit logging is a standard feature, which is useful if your compliance teams need to demonstrate log integrity rather than just log availability
Compliance posture
- PCI DSS
- SOC 2 Type 2
- ISO 27001
- FedRAMP Moderate
Pricing model
Credits-based. You purchase a pool of credits, and different activities consume them at different rates. SIEM ingest burns credits roughly eight times faster than standard log ingest, for example.
A regional multiplier applies: the US base rate is 1.0, but deploying in Frankfurt or Sydney adds 20%, and the European Sovereign Cloud adds 40%. That sovereignty premium is predictable, but it needs to be in the model from the start rather than discovered at renewal.
Regulatory fit
- DORA: The sovereign cloud option directly addresses EU data residency requirements, and the tamper-evident logging supports DORA’s evidence expectations
- MAS TRM: No blockers. The Tokyo and Sydney regions cover the APAC presence, and the security focus aligns well with MAS’s emphasis on continuous monitoring of critical systems
- FCA/PRA: Works well for financial institutions whose primary use case is security analytics with observability layered on
- APRA CPS 230: The Sydney region is available, but Sumo Logic does not hold IRAP
Best for
Teams for whom budgets are tight, and security is a high priority. Perhaps fledging FinTechs who need to demonstrate they’re secure more than they need to process high volumes of data.
Coralogix
Capabilities
Coralogix is built around a different premise to most observability tools on this list. Its Streama engine analyses telemetry data in-stream so alerting and dashboarding happen on the data as it flows through the pipeline, before it is indexed.
That means you can alert on 100% of your data while only indexing and paying for the fraction you need for interactive queries. For financial services environments generating large volumes of compliance and audit logs, that architecture directly addresses the cost problem that makes other observability platforms expensive at scale.
- Deployment: Multi-region SaaS. Not self-hosted, but the Compliance storage tier writes to your own S3 bucket, so your cold data sits in storage you control. That is not the same as full self-hosted deployment, but it is a stronger residency answer than most SaaS-only platforms offer
- Retention: Three tiers: Frequent Search (hot), Monitoring (warm) and Compliance (cold). The TCO Optimizer lets you route data to the appropriate tier per source, namespace or application. Whatever tier a log lands in, it is also written to your object storage, so long-term retention cost is driven by your S3 bill rather than by the platform’s pricing
Compliance posture
- PCI DSS
- SOC 2 Type 2
- ISO 27001
- FedRAMP Moderate (Coralogix U.S. GovOps)
Pricing model
Usage-based, per GB, with no per-user or per-host fees. All features are included on every paid plan, including 24/7 support. The effective per-GB rate varies by tier: Frequent Search is the most expensive, Compliance the cheapest. Because cold-tier data sits in your own S3, long-term retention cost is largely decoupled from the platform fee. For financial services teams required to retain years of audit data, that decoupling is where the business impact on your observability budget is smallest.
Regulatory fit
- DORA: The in-stream alerting architecture means detection does not depend on which retention tier the data lands in, which avoids the tiering trade-off that affects some other platforms. Your cold data sitting in your own S3 gives you direct control over long-term evidence
- MAS TrafRM: Should work well for detection and alerting. No APAC-specific sovereign deployment, but multi-region SaaS covers Singapore
- FCA/PRA: A good fit for UK-regulated FinTechs. The cost model makes it easier to justify retaining high-volume audit logs for the periods your compliance teams need without the bill becoming the dominant line item
- APRA CPS 230: No IRAP and no Australian region; weaker than the other options on this list.
Best for
Growth-stage FinTechs and mid-market financial institutions where observability cost is a constraint. A scaling PSP or neobank generating large volumes of transaction logs that need to be retained but not necessarily queried interactively.
How we can help
At Just After Midnight, we’ve helped APAC salary package firm GO Salary to resolve incidents in under 45 minutes on average and helped one of the UK’s biggest mobile finance providers to effectively monitor 10+ third-party services.
To find out how we could help you with observability, incident response or any of the other cloud and reliability issues we see our FinTech partners face across regions, just get in touch.
