DOE-driven material selection for power-dense aerospace gearboxes
Selecting the right gear material sits at the heart of every weight-conscious aerospace programme. For a power reduction gearbox destined for a geared turbofan, the choice between a case-hardening steel, a through-hardened stainless, or even a beryllium-aluminium composite can shift the gearbox mass by tens of kilograms, ripple through the whole engine pylon, and ultimately decide whether an airframer meets its block fuel budget on a long-haul sector from Sydney to Los Angeles. The OPTIMIZE Project treats this choice as something that should be solved statistically, not guessed, and that is where a structured design of experiments approach earns its place in the materials engineer’s workflow.
A DOE campaign lets you test multiple alloy candidates, heat treatments, and surface treatments under a single, statistically balanced testing matrix, then pull out the effects that actually matter for power density and mass. In a region like Australia, where research collaborations between CSIRO, the DSTG labs in Edinburgh (South Australia), and the aerospace cluster around Fishermans Bend in Melbourne are reshaping the local supply chain, this kind of rigorous methodology resonates with how engineers here already plan qualification campaigns. The aim of this write-up is to walk through how a DOE-driven selection process actually plays out, what data it produces, and how that data feeds into a defensible material call for a specific power and weight target.
Why material choice dictates the power-to-weight envelope
The power density of a gearbox is the product of three tightly coupled decisions: gear geometry, surface treatment, and base material stiffness. Of those three, the base material is the slowest to change and the most expensive to reverse, which is why it deserves the first hard look. A power reduction unit rated for, say, 8 MW of input power at 12 000 rpm needs gears that can carry Hertzian contact pressures well above 1.6 GPa while still being light enough to keep the centre-of-gravity envelope tidy for a low-slung engine installation.
Material density dominates the weight equation directly. Swapping a conventional through-hardened steel like AISI 4340 for a titanium alloy such as Ti-6Al-4V cuts density by roughly 45 percent, but the same swap can cut the allowable contact pressure by 25 to 30 percent unless the gear face width is increased. That tension between density and load-carrying capacity is exactly the kind of trade-off a DOE campaign is built to resolve, because it lets you hold geometry fixed in software while you sweep material parameters and surface treatment combinations.
Local context matters too. The Australian aerospace aftermarket is small compared with Toulouse or Seattle, but defence procurements through CASG often demand that any new alloy carry traceable pedigree from a mill listed on approved supplier registers. A DOE campaign doubles as a paper trail that shows the Commonwealth, in fair dinkum detail, why the chosen material is the safest bet for a flight-critical part.
Designing the experiment matrix
The first practical step is to define the response variables the design must satisfy. For a power and weight target, the two primary responses are usually specific power (kW per kg of gearbox) and a combined durability index that rolls in bending fatigue, rolling contact fatigue, and scuffing resistance. Secondary responses can include manufacturability ratings, cost per kilogram in Australian dollars, and a thermal compatibility score that reflects how the alloy behaves in the splash-lubricated environment of a reduction gearbox.
The factors in the matrix are then built around the candidate materials, the case depth after carburising or nitriding, and the surface finish after superfinishing. A typical full factorial at two levels across three materials and two case depths gives 12 runs per replication, which is a manageable workload for a single test rig. When interactions between alloy and case depth are suspected, and they usually are, a Taguchi L9 or a custom D-optimal design often produces sharper insights with fewer physical samples.
Engineers running campaigns out of the University of Melbourne’s precision-machining lab or the Advanced Manufacturing Growth Centre in Sydney have access to the kind of test rigs that can deliver rolling-sliding contact data at realistic pitch line speeds. With proper replication and randomisation, the resulting ANOVA output gives confidence intervals on each main factor, exposing whether a material effect is real or simply buried in test-to-test scatter. That level of rigour is what separates a defensible material recommendation from a back-of-the-envelope guess.
Candidate alloys and their baseline properties
Before any test rig spins up, the candidate list needs to be narrowed using density, elastic modulus, fatigue limit, and indicative cost. The table below captures a realistic shortlist for a power reduction gearbox aimed at the 8 MW class, with all values expressed in the units an Australian design office would normally pull from ASM data sheets and supplier datasheets.
| Material | Density (g/cm³) | Yield strength (MPa) | Bending fatigue limit (MPa) | Indicative cost (AUD/kg) | Suitability for 8 MW class |
|---|---|---|---|---|---|
| AISI 9310 (carburised) | 7.85 | 1180 | 620 | 18–24 | Strong baseline |
| M50 NiL (carburised) | 7.85 | 1240 | 680 | 32–40 | Strong for high speed |
| Pyrowear 53 | 7.80 | 1280 | 700 | 45–55 | Strong at temperature |
| Cronidur 30 | 7.65 | 1450 | 760 | 90–110 | Strong for corrosion |
| Ti-6Al-4V (nitrided) | 4.43 | 880 | 470 | 65–80 | Limited by load |
| AlBeMet 162 | 2.10 | 350 | 220 | 280–340 | Light but soft |
Reading across the rows, the steels cluster between 7.6 and 7.9 g/cm³, the titanium drops the density by almost half at the price of fatigue strength, and AlBeMet looks attractive on paper until the fatigue column is examined. None of these materials is disqualified at this stage, which is precisely why a DOE campaign is needed: the table shows plausible options, but it cannot tell the design team which combination of alloy, case depth, and surface finish delivers the best specific power at acceptable risk.
The same approach also helps filter out candidates that the supply chain cannot support. M50 NiL is a common import through Adelaide’s aerospace distributors, while Pyrowear 53 often requires longer lead times from European mills, a delay that can ripple across a programme schedule that already has to accommodate the long lead-in on a geared turbofan certification cycle.
Running the campaign and pulling out the winners
Once the matrix is locked, the physical testing begins. Power-circulating test rigs running back-to-back, often instrumented with torque transducers, vibration pickups, and oil-temperature probes, accumulate the rolling contact fatigue data that feeds the statistical model. In a well-planned DOE, each run also captures surface finish evolution, lubricant film thickness, and bulk oil temperature rise, because a material that looks brilliant at steady state can fall apart once thermal expansion starts to bite.
Analysis of the data typically starts with a Pareto chart of standardised effects. Alloys with the largest main effect on specific power move to the front of the ranking, and any two-factor interactions, for example alloy times case depth, that exceed the noise floor get a closer look. In several OPTIMIZE-style campaigns, the interaction term between alloy choice and superfinishing roughness has been the deciding factor, because a slightly rougher surface on a through-hardened steel can wipe out the fatigue advantage it carries on paper.
When the ANOVA is complete, the optimal settings can be predicted with confidence intervals, and a confirmation run at those settings is the final gate. Engineers at the DSTG laboratories near Salisbury routinely use this confirmation step to validate predictions against physical coupons, and the same workflow translates cleanly into a civil aerospace programme. The output is a single recommendation, expressed in plain language, that an engineering manager can sign off without needing to re-derive the statistics.
Translating DOE results into a flight-ready decision
A DOE result is not the end of the material selection journey; it is the start of the qualification journey. The chosen alloy now has to clear a sequence of gates: mill certification, forging trials, gear cutting trials on a CNC hobber, and then a full set of rig tests at the rated load spectrum. For an Australian programme, that often means engaging partners such as Quickstep at Bankstown Airport for composite tooling or the teams at RMIT’s aerospace precinct for additive-manufactured gear blanks, depending on where the supply chain decision lands.
Documentation matters as much as the numbers. The campaign needs a traceability matrix that links every factor in the DOE back to a design requirement, a test method, and an acceptance criterion. That matrix is what survives an audit by EASA, CASA, or the relevant defence customer, and it is what convinces a sceptical programme manager that the recommendation is more than a clever statistical trick. joint dynamic response is a separate but related concern that the same engineering team will need to address during the structural validation phase.
The objectives of the wider OPTIMIZE programme make this clear: the methodology is only useful when it feeds back into a real engineering decision on a real gearbox, and the material call is one of the most visible of those decisions. By the time the programme reaches its Critical Design Review, the DOE-driven material choice should sit in the trade study with full confidence intervals, a clear winner, and a credible fallback.
Practical guidelines for a DOE-led material study
A few habits tend to make the difference between a DOE campaign that delivers clear answers and one that fizzles into inconclusive scatter:
- Keep the response variables to two or three primary metrics; chasing too many outputs dilutes the statistical power.
- Choose factor levels that bracket realistic operating conditions rather than theoretical extremes.
- Replicate every run at least twice to separate alloy effect from rig noise, especially on smaller test rigs.
- Treat surface finish and lubrication as factors in the matrix, not as background constants.
- Use a confirmation run at the predicted optimum before publishing the recommendation.
- Document every step in a traceability matrix that an auditor can follow without phoning the test lab.
- Plan the supply chain assessment in parallel with the rig testing, because lead time can override fatigue performance.
If your team is staring down a power and weight target for an aerospace reduction gearbox and the material shortlist is more political than technical, the DOE route is worth the setup time. Start with a clean response definition, a tight factor list, and a rig that can run the matrix within a quarter, and the recommendation will look after itself. The OPTIMIZE team continues to publish the underlying methods and lessons on the project site, so engineers across the country can lift the workflow into their own programmes without reinventing the wheel.