Engineering hubs in Dehradun & Bengaluru · Delivering across 10 countries

nitesh@redcubical.com +91 90687 14658

REDCUBICALSYSTEMS

Services / Operate

Application support with targets we publish and report against

Tiered L1, L2 and L3 support with a published P1 to P4 severity matrix, coverage from business hours to genuine follow-the-sun, monitoring and on-call ownership, patch cadence, ITIL-aligned change and problem management, and verified backups with rehearsed recovery. Every target is measured and reported monthly, including the months we miss.

  • P1 response in 15 minutes, with updates every 30 minutes until resolved
  • Runbooks we author and own, not documentation you have to maintain
  • Problem management that closes recurring incidents permanently
  • A structured 6 to 10 week transition before we accept the pager

At a glance

Tiers
L1 triage, L2 diagnosis, L3 engineering
Coverage models
Business hours, extended 16 hours, follow-the-sun 24x7
P1 response
15 minutes, updates every 30 minutes
Availability target
99.95% on supported architectures
Client overlap hours
4–8 hrs daily across 10 markets
Transition period
6 to 10 weeks before we take the pager
Onboarding to first ticket
10 business days on a documented system

Service definition

Support tiers and the severity matrix

Tiers and severity

L1 resolves documented issues from a runbook and triages everything else. L2 diagnoses undocumented problems with logs, traces and queries, and applies configuration fixes. L3 changes code, schema and architecture. Severity is assigned on business impact, not on how urgent the requester feels, and the matrix is published so classification is not a negotiation during an incident.

Support tier definitions
TierWhat it resolvesTypical share of ticketsSkills and accessEscalates when
L1 — service desk and triageAccess and permission requests, account and authentication issues, data corrections with a documented procedure, report reruns, known issues with a published workaround, and classification and routing of everything elseRoughly 55 percent once the knowledge base has maturedRunbook execution, read access to logs and dashboards, tightly scoped write actions through approved tooling onlyNo runbook exists, the runbook did not resolve it, or the issue is classified P1 or P2
L2 — application support engineeringUndocumented issues diagnosed from logs, traces and database queries. Configuration changes, cache and queue interventions, data repair under change control, integration failure diagnosis, and third-party escalationRoughly 30 percentRead access to production data with audit, write access under change control, full observability tooling, and vendor support portal accessA code, schema or infrastructure change is required, or root cause is in application logic
L3 — engineering and root causeCode defects, schema changes, performance engineering, architectural fixes, security remediation, and permanent resolution of recurring problems raised through problem managementRoughly 15 percentFull development environment, pipeline access, deployment authority under change management, and ownership of the codebaseRarely. L3 escalates to the vendor for platform defects, or to you for commercial and product decisions
Service delivery managementNot a resolution tier. Owns the service relationship, SLA reporting, the monthly review, escalation ownership, continual improvement and the problem management backlogNot ticket-basedNamed individual, not a rota. Attends your governance forumEscalates commercially to your account sponsor and ours

The share of tickets by tier is the metric that shows whether the service is improving. A mature engagement pushes work downward: today L2 diagnoses an issue, next week it is a runbook and L1 resolves it in four minutes. If your L3 percentage is not falling over the first two quarters, either the knowledge base is not being maintained or there is an underlying problem that incident management keeps treating instead of solving.

Severity matrix with response, resolution and escalation
SeverityDefinitionResponse targetResolution targetUpdate cadenceEscalation path
P1 — CriticalComplete outage, or a critical business function unavailable, or confirmed data loss or a security incident in progress. No workaround exists. Revenue or safety impact is immediate15 minutes, 24x7 on applicable coverage4 hours to service restorationEvery 30 minutes, including when there is nothing newIncident commander at declaration, engineering lead at 1 hour, service delivery manager and your sponsor at 2 hours, executive at 4 hours
P2 — HighMajor function degraded or unavailable with a workaround that is painful but viable. A significant subset of users affected. Performance degraded to the point of interfering with work1 hour within coverage hours8 business hoursEvery 2 hoursL2 lead at 2 hours, engineering lead at 4 hours, service delivery manager at 8 hours
P3 — MediumA non-critical function is not working correctly, or a minor function is unavailable. A workaround exists and is acceptable. Limited user impact4 business hours3 business daysDailyL2 lead at 1 business day, service delivery manager at 3 business days
P4 — LowCosmetic defect, documentation error, enhancement request, or a question. No functional impactNext business dayNext planned release cycleWeekly, or on status changeReviewed in the monthly service review rather than escalated

Two things we are deliberate about. Response is a contractual commitment and is measured to the minute. Resolution is a target, because a defect requiring a third-party fix or a change that cannot be made safely in four hours will not close in four hours, and promising otherwise would be dishonest. Second, severity is set on business impact against the published definitions. If we disagree with your classification we escalate rather than silently downgrade, and the audit trail shows both positions.

Commercials

Coverage models, indicative pricing, and what is out of scope

Coverage and pricing

Three coverage models. Business hours in one region, extended at roughly 16 hours across two shifts, and follow-the-sun 24x7 with named on-call. True 24x7 costs three to four times business hours rather than three times the hours, because it needs enough engineers for a sustainable rotation, not the same people paged at night.

Coverage models and indicative monthly pricing
ModelHours coveredP1 responseOut-of-hours handlingIndicative monthly (USD)Right for
Business hours, single region9 hours a day, 5 days a week, aligned to your timezone within our overlap15 minutes within hoursVoicemail and email queued to the next business morning. No on-call3,500 to 8,000Internal tools, back-office systems, and anything where an overnight outage is inconvenient rather than costly
Extended, two shiftsRoughly 16 hours a day, 5 to 6 days a week, covering your working day plus evening15 minutes within hours, 1 hour outsideBest-effort on-call for P1 only, with a defined but longer response target7,000 to 15,000Customer-facing applications with concentrated usage hours, and multi-region users across two or three timezones
Follow-the-sun 24x724 hours, 7 days, 365 days, with handover between shifts15 minutes at any hour, including weekends and public holidaysThere is no out of hours. A named engineer is on duty at all times with a documented handover15,000 to 35,000Revenue-critical platforms, regulated services, healthcare, payments, and anything with a contractual uptime commitment to your own customers

Prices are indicative bands for scoping and move with the number of systems, ticket volume, technology breadth, the reserved L3 engineering capacity you want each month, and whether data residency constraints limit where engineers can be located. A firm proposal follows a discovery call and a review of your last six months of ticket history, because volume and mix drive cost far more than system count does.

In scope

  • Incident management across all severities, including out-of-hours P1 on applicable coverage.
  • Service requests from a defined catalogue: access, permissions, data corrections, report reruns, configuration changes.
  • Monitoring and alert ownership. We receive the alerts, we act on them, and we tune them so the volume stays meaningful.
  • Bug fixes in the supported codebase, sized within the agreed monthly L3 allocation.
  • Security and dependency patching to the published cadence, including emergency patching outside a window when severity requires it.
  • Backup verification and restore testing to a schedule, with results reported rather than assumed.
  • Disaster recovery exercises at agreed intervals, measured against RPO and RTO targets.
  • Runbook and knowledge base authoring and maintenance, owned by us and readable by you.
  • Change management for supported systems, including release execution.
  • Problem management to permanently close recurring incidents.
  • Monthly service review with a named service delivery manager.

Explicitly out of scope

  • New feature development. Quoted as a project. A support contract used as a development team ends badly for both sides.
  • Major version upgrades and migrations. Framework major versions, database engine upgrades, cloud migrations. Planned work with its own budget and rollback plan.
  • Third-party product support where the vendor is the correct escalation route. We will raise and chase the ticket. We cannot fix their code.
  • End-user device and desktop support, laptop imaging, printers and office networking, unless separately contracted.
  • Business process work. Data entry, manual reconciliation, invoice processing and report interpretation are operational activities, not application support.
  • Anything requiring a licence, certification or regulated status we do not hold, including audit opinions and legal or clinical judgement.
  • Systems outside the agreed inventory. New systems are added by a documented change with a transition period, not by being mentioned in a ticket.
  • Unlimited L3 engineering. A monthly allocation in engineering days, with overflow quoted or carried by agreement.

How it runs

Runbooks, monitoring handover, patching and change management

Operational practice

We author and own the runbooks, take ownership of the alert stream rather than merely receiving it, patch to a published cadence, and run every change through a classification: standard, normal or emergency. Standard changes are pre-approved and low-risk. Normal changes need review. Emergency changes are permitted, and reviewed retrospectively without exception.

Runbook and knowledge base ownership

  • We write them, we own them. Support documentation is a deliverable of the service, not homework for your team.
  • One runbook per alert. If an alert has no runbook, either we write one or we delete the alert. An alert nobody knows how to action trains people to ignore alerts generally.
  • Structure that works at 3am. Symptom, verification steps, immediate mitigation, diagnosis path, escalation trigger, and rollback. Written for someone tired and under pressure.
  • Tested by execution. Every runbook is executed by an engineer who did not write it, in a non-production environment. Anything they cannot follow is rewritten. This is where most inherited documentation fails.
  • Updated at the point of use. Every incident either follows a runbook or produces one. A runbook untouched for a year on a system that changes monthly is fiction.
  • Kept in your systems. Confluence, Notion or your repository. You retain everything at the end of the contract, which matters on the day you change supplier.

Monitoring and alerting handover

  • Alert inventory first. Every existing alert reviewed for whether it is actionable, correctly targeted and correctly thresholded. In most takeovers a third are deleted immediately.
  • Symptom-based alerting. We alert on user-visible impact and error budget burn rate, not on every CPU threshold. Cause-based alerts belong on dashboards, not in a pager.
  • Coverage gaps closed. Silent failures are the dangerous class: a queue not draining, a scheduled job that did not run, a file that did not arrive, a certificate approaching expiry. We alert on absence, not only on error.
  • Routing to the right tier. An alert with a runbook goes to L1. An alert requiring judgement goes to L2. Only genuine P1 signals reach out-of-hours on-call.
  • Noise budget. More than two pages per on-call shift triggers a review of the alerting, not of the engineer.
  • Your visibility is preserved. You keep access to every dashboard and alert. We are not a black box, and you should be able to see what we see.
Patch and dependency management cadence
CategoryCadenceWindowDowntimeRollback
Critical security patches (CVSS 9.0+ or actively exploited)Within 72 hours of confirmed triageEmergency window, outside standard change approval with retrospective reviewUsually zero via rolling replacement. Where unavoidable, agreed with you before proceedingPrevious image or package version retained and pre-tested
High severity security patches (CVSS 7.0 to 8.9)Within 14 daysNext scheduled maintenance windowZero downtime expectedStandard rollback via previous artefact
Operating system and runtime patchingMonthlyAgreed monthly window, typically weekend early morning in your timezoneZero on multi-instance architectures. Single-instance systems need a windowImmutable infrastructure means rollback is redeploying the previous image
Application dependency updatesContinuous via automated pull requests, merged weeklyNormal release processNoneStandard deployment rollback. The test suite is the gate
Database minor version and engine patchesQuarterly, aligned to the provider maintenance scheduleAgreed window with at least 10 business days noticeTypically 1 to 5 minutes on a multi-AZ failover. Longer for single-instancePoint-in-time recovery available. Snapshot taken immediately before
Major version upgradesPlanned project, not patchingIts own plan, testing cycle and rollback designDepends entirely on the upgradeDocumented per project, rehearsed before execution

Patching is where support contracts quietly fail. It is invisible when done and catastrophic when neglected, so it is easy to defer under delivery pressure. We report patch currency every month with the count and age of outstanding items by severity, which makes deferral a visible decision rather than a silent drift.

Change management

  • Standard changes. Pre-approved, low-risk, repeatable and well-understood: a routine dependency update, a scaling adjustment within agreed bounds, a documented configuration change. No individual approval needed, logged automatically, reviewed in aggregate monthly.
  • Normal changes. Everything else. A written change record with risk assessment, test evidence, rollback plan and target window. Approved by the change authority appropriate to the risk level.
  • Emergency changes. Permitted during a P1 when the alternative is prolonged outage. Verbal approval from the incident commander and your on-call contact, then a full change record and retrospective review within two business days. No exceptions to the retrospective, because emergency change is where controls quietly erode.
  • Change advisory board. A fortnightly forum for normal changes above a risk threshold, with your technical representative present. Reviews the forward schedule, conflicts between changes, and business timing such as month-end or a peak trading period.
  • Change freeze periods agreed in advance around your critical business dates, with a documented exception route for security patching.
  • Change success rate reported monthly. A failed change is one that was rolled back or caused an incident. Below 95 percent means the process needs attention, not that the engineers do.

Problem management, ITIL-aligned

  • Incidents and problems are different things. An incident is restoring service now. A problem is the underlying cause producing repeated incidents. Conflating them is why the same outage happens quarterly for three years.
  • A problem record is raised when the same incident category occurs three times in 90 days, after every P1 regardless of recurrence, or when an incident review identifies a systemic cause.
  • Root cause analysis using five whys or fault tree analysis, documented, blameless, and reaching past the proximate trigger to the condition that allowed it.
  • Known error database. Problems with an identified cause but no economic fix yet, recorded with their workaround so L1 can resolve related tickets immediately.
  • Problems carry an owner and a target date, and are reviewed monthly with you. A problem backlog nobody reviews is just a list.
  • Recurring incident analysis is a standing item in the monthly review. Top five problem categories by ticket volume, with what is being done about each.
  • The commercial logic. Every permanently closed problem removes tickets from the queue. Over a year this is what shifts an engagement from firefighting to genuine stability, and it is the part of the service most easily neglected under ticket pressure.

Assurance

Backup verification, DR testing and the monthly service review

Backup and recovery assurance

A backup that has never been restored is a hypothesis. We verify restores on a schedule, run disaster recovery exercises against stated RPO and RTO targets, and report the measured recovery time rather than the designed one. When the measured figure misses the target, that goes in the report and becomes a problem record.

Backup, recovery and DR testing targets
System classRPO targetRTO targetBackup regimeVerificationDR exercise
Tier 1 — revenue criticalUnder 5 minutesUnder 1 hourContinuous replication plus point-in-time recovery, cross-region copies, immutable retention for 35 daysAutomated restore to an isolated environment weekly, with an integrity check on the restored dataFull failover exercise twice a year, including a region failure scenario
Tier 2 — business importantUnder 1 hourUnder 4 hoursHourly incremental, daily full, cross-region copy, 30-day retentionAutomated restore monthlyAnnual failover exercise
Tier 3 — internal and supportingUnder 24 hoursUnder 24 hoursDaily snapshot, 14 to 30-day retentionQuarterly restore testAnnual tabletop exercise rather than a live failover
Object and document storageNear zero with versioning enabledUnder 4 hours for a bulk restoreVersioning plus cross-region replication, lifecycle policy to archive tiers, object lock where compliance requires immutabilityQuarterly sample restore including from the archive tier, because archive retrieval time is what surprises peopleIncluded in the tier failover exercise
Secrets and encryption keysZeroUnder 1 hourManaged store with automatic backup, plus documented escrow for the keys that would prevent recovery if lostRecovery procedure tested twice a yearIncluded in the DR exercise, because an untested key recovery path makes every other backup worthless

Two things we insist on measuring. First, the actual recovery time in the exercise, not the designed one. Restores are routinely two to five times slower than assumed, particularly from archive tiers where retrieval latency is measured in hours. Second, whether the restored data is correct rather than merely present. A restore that completes and produces a corrupt or partial dataset counts as a failed test in our reporting.

What the monthly service review covers

  1. SLA attainment per severity: response and resolution, met and missed, with every miss explained individually rather than absorbed into a percentage.
  2. Ticket volume by category and by system, with the trend over six months. Rising volume in one category is the signal that starts a problem record.
  3. Tier distribution. The L1, L2 and L3 split, showing whether knowledge is genuinely transferring downward.
  4. MTTR by severity, with the distribution rather than only the mean, because one nine-hour P2 hidden behind a good average is the thing worth discussing.
  5. Recurring problem analysis. Top five problem categories, root cause status, owner and target date for each.
  6. Availability against target, with the measurement basis stated and every period of degradation listed.
  7. Change record. Changes delivered, success rate, failed changes with cause, and the forward schedule.
  8. Patch currency by severity, with the age of the oldest outstanding item.
  9. Backup and DR evidence. Restore tests performed, measured recovery times, and any exercise that missed its target.
  10. Capacity and cost trend, including anything heading towards a limit.
  11. Continual improvement. What we changed last month, what we propose next, and what we need from you.

Taking over an existing system

  1. Weeks 1 to 2: discovery and access. System inventory, architecture and data flow mapping, dependency and integration list, access provisioning through your process, and a review of the last six months of tickets to understand real volume and mix.
  2. Weeks 2 to 4: knowledge transfer. Recorded sessions with the incumbent team or the current owner. Where documentation does not exist, we read the code and write what is missing. This is normal and we budget for it.
  3. Weeks 3 to 6: runbook authoring. The top twenty scenarios by historical frequency and impact, each executed in a non-production environment by someone who did not write it.
  4. Weeks 4 to 6: monitoring review. Alert inventory, noise removal, gap closure, and routing configured to the correct tier.
  5. Weeks 5 to 8: shadow period. The incumbent leads, we observe and respond in parallel without authority. Every gap this exposes is cheaper to close now than after cutover.
  6. Weeks 7 to 9: reverse shadow. We lead, the incumbent observes and can intervene. This is the point at which most residual documentation gaps surface.
  7. Week 9 to 10: cutover. Formal handover with an agreed rollback to the previous arrangement, a hypercare period of two weeks at elevated staffing, and daily reviews during hypercare.
  8. Why we will not compress this. A same-week start on an undocumented system means the first P1 is handled by people reading the code for the first time. We have declined engagements on that basis, and we would rather lose the contract than run the service badly.

Honest answer

What an SLA guarantees, and why service credits are not a remedy

What SLAs actually mean

An SLA guarantees a response time, an update cadence, an escalation path and a measured, reported outcome. It does not guarantee a fix within a window, and it does not guarantee the absence of outages. Service credits, typically 10 to 25 percent of one month fee, are an accountability mechanism, not compensation for what an outage costs you.

What an SLA does and does not cover
ClauseWhat it genuinely guaranteesWhat it does not
Response timeA qualified engineer acknowledges and begins work within the stated period, measured from ticket receipt and correct classification. This is a hard commitment and we report against it to the minuteThat the problem is understood, diagnosed or solved at that moment. Response is the start of work, not the end
Resolution targetContinuous effort at the stated priority, with the update cadence maintained and escalation triggered on scheduleA fix by the deadline. A third-party defect, a data corruption requiring careful reconstruction, or a change that cannot be made safely at speed will exceed it
Availability percentageA measured figure against a stated basis, reported monthly with every period of unavailability listedWhat most people assume. Read the exclusions: agreed maintenance, third-party provider failure, your own changes, and force majeure are typically excluded. Whether degraded counts as unavailable is a definitional question that changes the number materially
Escalation pathNamed individuals at each level with contact details, reachable within the stated timeframe, updated when people changeThat escalation makes the fix arrive faster. It ensures the right people know and can make decisions, which is not the same thing
Service creditsA financial acknowledgement of a miss, and a governance trigger that forces a reviewCompensation. A credit worth a fraction of one month fee will not approach the cost of a serious outage. Consequential loss is almost always excluded, in our contracts and in every supplier contract you will be offered
Coverage hoursA named engineer on duty and reachable during the stated hours, with documented shift handoverThat the specific person who knows your system best is the one who answers. That is what a dedicated pod buys, at a higher price

We publish this table because the most damaging thing in a support relationship is a client who believed the SLA meant something it did not, discovering the difference during their worst week. Read your exclusions before you sign, ours included.

  • Ask what the availability measurement basis is and what counts as excluded. The same architecture can be reported at 99.5 or 99.95 percent depending on definitions
  • Ask when the last restore test was run and what the measured recovery time was. If the answer is a designed figure rather than a measured one, there has not been a test
  • Ask who is on call tonight by name, and what they would do first for your most likely failure. A vague answer tells you the runbooks are thin
  • Ask for the last three months of SLA reports including misses. A supplier with no misses in a year is either not measuring or not reporting honestly
  • Ask how many problem records were permanently closed last quarter. This distinguishes a support team from a ticket queue
  • Ask what happens to your runbooks and knowledge base at contract end. If they live in the supplier tooling, you are paying to build an exit barrier against yourself

Answers

Managed IT and application support questions

What do L1, L2 and L3 support actually resolve?

L1 handles known and documented issues from a runbook: access and permission requests, password and account problems, data corrections with an existing procedure, and triage plus classification of everything else. L2 diagnoses undocumented problems using logs, traces and database queries, and applies configuration changes and known fixes. L3 is engineering: code changes, schema changes, architectural fixes and permanent resolution of recurring problems. Roughly 55 percent of tickets close at L1, 30 percent at L2 and 15 percent reach L3 once a knowledge base has matured.

What does 24x7 support cost?

Indicatively, business hours single-region coverage runs 3,500 to 8,000 US dollars per month, extended 16-hour coverage 7,000 to 15,000, and genuine follow-the-sun 24x7 with named on-call and a 15-minute P1 response 15,000 to 35,000. The range moves with system count, ticket volume, technology breadth and how much L3 engineering capacity is reserved. True 24x7 requires enough engineers for a sustainable rotation, which is why it costs roughly three to four times business hours rather than three times the hours.

What are your P1 response and resolution targets?

P1, meaning a complete outage or critical business function unavailable with no workaround: 15-minute response, 4-hour resolution target, updates every 30 minutes. P2: 1-hour response, 8-hour target, updates every 2 hours. P3: 4-hour response, 3-business-day target, daily updates. P4: next business day response, next release cycle target. Response is a commitment and is measured. Resolution is a target, because some defects are genuinely not fixable within a fixed window.

What is explicitly out of scope?

New feature development, redesigns, major version upgrades and migrations, third-party product support where the vendor is the correct escalation route, end-user device and desktop support unless separately contracted, business process and data entry work, and anything requiring a licence or certification we do not hold. Out-of-scope work is quoted separately as a project. We keep the boundary explicit because scope creep is how a support contract quietly becomes an understaffed development team.

How do you take over support of a system you did not build?

A structured six to ten week transition. Discovery and access, knowledge transfer sessions with the incumbent team or the documentation that exists, runbook authoring for the top twenty scenarios, monitoring and alert review, a shadow period where we observe and the incumbent leads, then a reverse-shadow period where we lead and they observe, then cutover with an agreed rollback. We do not accept a support contract on an undocumented system with a same-week start date, because the first P1 would be handled by people reading the code for the first time.

How often do you patch, and does it require downtime?

Security patches for critical CVEs within 72 hours of triage, high within 14 days. Operating system and runtime patching monthly in a standard maintenance window. Dependency updates continuously via automated pull requests with the test suite as the gate. Major version upgrades are planned projects, not patches. Most patching is zero-downtime on a properly architected system through rolling replacement. Where downtime is unavoidable we agree the window at least ten business days ahead.

What do service credits actually get us?

Very little, honestly. A typical credit of 10 to 25 percent of one month fee for an SLA miss is a small fraction of what a serious outage costs your business. Credits are an accountability signal and a governance trigger, not compensation. What genuinely protects you is architecture that fails gracefully, tested backups, rehearsed recovery, and a supplier whose commercial interest is in prevention. We include credits because you should be able to hold us to something. We tell you plainly that they are not a remedy that makes you whole.

What does an SLA guarantee, and what does it not?

It guarantees a response time to a raised and correctly classified ticket, an update cadence, an escalation path with named people, and a measured and reported outcome. It does not guarantee a fix within a window, because some defects require third-party action or a change nobody can safely make in four hours. It does not guarantee the absence of outages. And an availability figure like 99.9 percent is a measurement basis you should read carefully, because what counts as excluded maintenance and what counts as degraded determine what the number means.

Start with a support readiness assessment

Two weeks, fixed fee. You get a system inventory with a support risk rating, six months of ticket history analysed by category and tier, a runbook gap list, an alerting review, a backup and DR verification check, and a coverage and pricing recommendation. Yours to keep whether or not you appoint us.