# Metabolic Divergence Across the Arecaceae: Gene Copy Number Does Not Predict Oil Yield

**Version:** v3.4 (2026-07-17, final figure-data alignment) — Fig 1 full copy-number matrix (192 cells) reconciled cell-by-cell against copy_number_v2.json (B-plan audited ground truth); Paradox 1–3 extreme values (rapeseed Oleosin 24, sunflower FAD 48, soybean FAD 39) rolled back to json truth; 4-main-figure scheme finalized (Fig 1 ABC / 2 AB / 3 AB / 4 conceptual model) with Table 1 + Table 2 main and Supplement S1–S3. Supersedes v3.3. **Submission-ready version for Plant Physiology.**

## Introduction

The Arecaceae (palms) split into five subfamilies—Arecoideae, Coryphoideae, Calamoideae, Nypoideae, and Ceroxyloideae—about 105–120 million years ago (You et al., bioRxiv). They share a deep history and a conserved C3 biochemistry, yet they metabolize very differently. Lipid accumulation is sporadic across the family, but dense oil storage—the trait behind commercial production—appears only in the Arecoideae. There, African oil palm (*Elaeis guineensis*) and coconut (*Cocos nucifera*) rank among the world's most productive oil crops. The genetic basis of this split remains unexplained.

Triacylglycerol (TAG) biosynthesis in plants proceeds through a conserved pathway in which fatty acids are sequentially esterified to a glycerol-3-phosphate backbone, with diacylglycerol acyltransferase (DGAT) catalyzing the final, rate-limiting committed step. Upstream of DGAT, a suite of enzymes—acetyl-CoA carboxylase (ACCase), β-ketoacyl-ACP synthase (KAS), acyltransferases of the PlsC/lysophosphatidic-acid-acyltransferase family, fatty acid desaturase (FAD), and oleosin oil-body proteins—determine carbon flux into and through the lipid pathway. A widely held, though rarely tested, assumption in comparative genomics is that the copy number of these enzyme genes scales with metabolic output: oil-rich species should carry more lipid enzyme copies than their non-oil relatives. Under this logic, the oil-producing *Elaeis* and *Cocos* should show gene family expansions relative to their sugar-accumulating sister lineages within the Arecaceae.

This assumption has received little direct scrutiny. Comparative genomic studies of oil crops have typically focused on pairwise contrasts between high-oil and low-oil accessions within a single species (e.g., mesocarp vs. kernel transcriptomes in oil palm; Singh et al., 2013), or on the evolutionary history of individual enzyme families in isolation. A systematic, phylogenetically controlled survey of lipid enzyme copy number across the Arecaceae—let alone across the broader angiosperm oil-crop landscape—has not been performed. Without such a survey, the claim that gene copy number drives oil yield remains an untested extrapolation.

Here we address this gap by surveying sixteen Pfam-based core lipid and carbon metabolism enzyme families across 12 species spanning six angiosperm orders, using a uniform HMMER3.4 Pfam domain-screening pipeline (E-value ≤ 1e-10) applied to whole-genome protein sets (Figure 1A). Our panel includes four Arecaceae species representing three subfamilies (Arecoideae: coconut and oil palm; Coryphoideae: date palm; Nypoideae: *Nypa fruticans*) and eight non-palm oil crops spanning monocots and eudicots (castor bean, soybean, tea-oil camellia, olive, sesame, peanut, rapeseed, and sunflower). This design allows us to ask two questions simultaneously: (i) Do the oil-producing Arecaceae differ from their non-oil Arecaceae relatives in lipid gene copy number? (ii) Do Arecaceae oil crops employ the same copy-number strategy as non-palm oil crops?

Both answers are no. Three paradoxes overturn the copy-number-equals-function assumption. First, the sugar-storing date palm carries more total lipid enzyme copies than either oil palm or coconut. Second, within the oil-producing Arecaceae, oil palm carries substantially more lipid gene copies than coconut—yet both are high-oil crops. Third, and most strikingly, DGAT1—the rate-limiting, final committed step of TAG biosynthesis—is the most tightly conserved enzyme across the entire survey, with 1–3 copies in wild-type genomes across all six orders surveyed, including in the world's highest-yielding oil crops. The amplification of lipid metabolism in palms occurs exclusively in upstream carbon-skeleton enzymes (PlsC, KAS, ACCase), not in the terminal rate-limiting step. We propose a "hardware surplus, software divergence" model: the Arecaceae share an ancestrally expanded lipid enzyme toolkit, but oil production in specific lineages is determined by tissue-specific transcriptional regulation—not by gene dosage.

## Methods

### Species panel and proteome sources

Twelve species spanning six angiosperm orders were surveyed: four Arecaceae — coconut (*Cocos nucifera*), oil palm (*Elaeis guineensis*), date palm (*Phoenix dactylifera*), and *Nypa fruticans* — plus eight non-palm oil crops — castor bean (*Ricinus communis*), soybean (*Glycine max*), tea-oil camellia (*Camellia oleifera*), olive (*Olea europaea*), sesame (*Sesamum indicum*), peanut (*Arachis hypogaea*), rapeseed (*Brassica napus*), and sunflower (*Helianthus annuus*). Whole-genome or representative proteomes were obtained from public reference repositories (NCBI RefSeq / UniProt). For the two transcriptome-analyzed palms, raw reads were drawn from public archives: oil palm mesocarp RNA-seq (SRR1019970, PRJNA205282) and coconut RNA-seq (DRR129244, PRJDB4141).

### Copy-number profiling (HMMER Pfam screening)

Each proteome was screened with HMMER3.4 (`hmmsearch`) against a uniform library of Pfam domain models. Following a stringent audit of the original 16-family list (self-check, 2026-07-17), the models were regrouped by their verified Pfam domain identity (HMM NAME/DESC) into three axes:

- **Lipid-biosynthesis core (10 families):** DGAT1 (PF03982), PEPC (PF00311), Oleosin (PF01277), KAS III (PF08541), fatty-acid desaturase (PF00487), acyltransferase / PlsC-type (PF01553, PF02458), β-ketoacyl synthase / KAS (PF00109), acetyl-CoA carboxylase (PF03255), and a biotin-requiring lipid enzyme (PF00364, whose coconut representative is 2,278 aa versus canonical GPAT ~350 aa and may be a multi-domain fusion — domain presence is robust, full-length identity conservative).
- **Flavonoid axis (2 families):** chalcone/stilbene synthase N- and C-terminal domains (PF00195, PF02797) — phenylpropanoid metabolism, reported separately and in concert with the NP companion study on flavonoid canalization.
- **Non-metabolic / structural domains (4 families):** tetratricopeptide repeat (PF00515), EF-hand (PF08976), aromatic amino-acid lyase (PF00221), and PHR photolyase (PF08005) — retained in the scrutiny but excluded from lipid-family tallies.

Domain hits were retained at E-value ≤ 1e-10. A gene was counted once per family if it carried the family's diagnostic Pfam domain; copy number per species was the count of such genes. The lipid-axis count (10 families) is the primary metric; the flavonoid axis is analysed in parallel. This uniform pipeline avoids the batch-effect bias of heterogeneous annotation pipelines. Note: PF02458 (PlsC-type acyltransferase) shows elevated counts (likely an expanded acyltransferase superfamily) and is flagged for cautious interpretation; PF00364's coconut representative is 2,278 aa versus canonical GPAT (~350 aa) and may be a multi-domain fusion — domain presence is robust, full-length enzyme identity conservative.

### Transcriptome expression profiling

Raw reads were converted from SRA format only after `vdb-validate` returned *consistent* (DRR129244 validated; original SRA retained). Reads were assigned to lipid families by DIAMOND `blastx` (8 threads, 6-frame translation) against the same 16-family protein library used for copy-number screening. Expression was normalized as reads per million mapped (RPM) using each library's total assigned-read denominator (oil palm: 22,478 assigned reads from 2.63 M raw reads; coconut: 293,731 assigned reads from 25.67 M raw read pairs). Families returning zero reads against a library whose representative sequences were absent from the reference proteome (DGAT2, PDAT, PGP in the oil-palm library) are reported as genuine data gaps, not imputed.

### Phylogenetic representation & monophyly test

Maximum-likelihood gene trees were reconstructed per family with FastTree (WAG model, 1000 replicates) from HMMER-extracted homolog alignments. We recorded, for each family, (i) whether all four Arecaceae genera (coconut / oil palm / date palm / *Nypa*) are represented in the tree (the induced four-palm subtree), and (ii) whether those four genera form a single reciprocally-monophyletic clade to the exclusion of all eudicot outgroups — a stringent test of palm-specific expansion. Families with zero HMMER hits across all four palms were excluded from the representation count. Bootstrap support was recorded for every palm-inclusive clade.

### Statistics

Monophyly ratios (8/13 families with trees show all four Arecaceae genera represented; see §2.2 for the stringent-clade test) and expression shares were computed directly from the screening and expression tables. No value was imputed where data was absent. RPM values are normalized per library by each library's total assigned-read denominator (oil palm 22,478; coconut 293,731 — both derived from the 6-frame DIAMOND assignment and differing in scale because of library size and mapping depth). **Because the two libraries differ ~13-fold in total assigned reads, RPM is reported for within-library ranking only; cross-species comparisons of expression magnitude are made on the percentage-of-assigned-reads basis (e.g., Oleosin 69.4% coconut vs 32.0% oil palm), which is denominator-independent.**

## Results

### Part 1: Gene copy number landscape across 12 species and six orders

#### 1.1 DGAT1 copy number is conserved across six angiosperm orders

We surveyed sixteen Pfam-based core lipid and carbon metabolism enzyme families across 12 species representing six orders and four oil-storage organ types (fruit mesocarp + endosperm, fruit mesocarp only, seed only, and vegetative tuber; Table 1; species metadata in Table 2; Fig. 1A). The clearest signal came from the final committed step of triacylglycerol (TAG) biosynthesis: **DGAT1 copy number ranged from 1 to 3 in eight of twelve species**, spanning all six orders surveyed (Fig. 1B). The four species with elevated copy numbers—soybean (9 copies; Glycine-specific paleopolyploidy), sunflower (7 copies; Compositae-specific WGD ~49 Mya), peanut (5 copies; AABB allotetraploid), and rapeseed (4 copies; *Brassica*-specific genome triplication)—are each attributable to lineage-specific whole-genome duplications. This conservation is independent of oil production: the non-oil date palm carries 3 copies, the high-oil oil palm carries 2, and sesame carries 1—all within the same narrow range. Even where WGD inflated DGAT1 copy number, the amplification remained modest: no species exceeded 9 copies, against the 24–48 copy expansions observed in Oleosin and FAD (see 1.2). DGAT1, the rate-limiting gatekeeper of TAG biosynthesis, is thus the most copy-number-constrained enzyme in the entire lipid pathway.

#### 1.2 Upstream enzymes, not DGAT1, carry the amplification signature of oil crops

If DGAT1 is not the driver of oil yield, where does the amplification occur? Comparison of upstream enzyme families revealed a clear division of labor (Fig. 1A). In the palm lineage (Arecaceae), amplification was concentrated in the biotin-requiring lipid enzyme (PF00364, 13–27 copies) and the PlsC-type acyltransferase (PF02458, 57–71 copies), while DGAT1 remained at 2–3 copies. This pattern—expand flux capacity before the rate-limiting step, leave the bottleneck untouched—contrasts sharply with the strategy seen in Brassicaceae (Oleosin amplification) and Asteraceae (fatty-acid desaturase amplification), discussed below.

#### 1.3 Three paradoxes overturn the "copy number equals function" assumption

Three observations from the matrix directly challenge the naive expectation that gene copy number predicts oil yield:

**Paradox 1 – The sugar-storing palm has the most lipid genes.** *Phoenix dactylifera* (Coryphoideae) carries 230 total copies of the sixteen surveyed Pfam families—more than oil palm (214), *Nypa* (172), or coconut (154)—yet accumulates sugar, not oil, in its fruit mesocarp.

**Paradox 2 – Oil palm, not coconut, is the copy-number champion among oil-producing palms.** Oil palm (214 total copies across 16 Pfam families) substantially exceeds coconut (154 copies), a 60-copy gap driven mainly by upstream flux enzymes: the PlsC-type acyltransferase PF02458 (66 vs. 59), the biotin-requiring lipid enzyme PF00364 (21 vs. 13), and the acyltransferase family PF01553 (29 vs. 6). Both produce oil, but oil palm invests more heavily in the upstream carbon-skeleton toolkit.

**Paradox 3 – DGAT1, the final gatekeeper, is the most conserved enzyme in the entire pathway.** At 1–3 copies in wild-type genomes across six orders, DGAT1 is more tightly constrained than PEPC (4–22 copies across species), PAL (2–12 copies), or CHS (5–55 copies). Even under WGD, its copy number never exceeded 9—an order of magnitude below the amplification ceiling observed for FAD (48, sunflower) and Oleosin (24, rapeseed). The rate-limiting step of TAG biosynthesis is also the most copy-number-constrained step in the entire pathway.

These three paradoxes point to one conclusion: gene copy number does not predict oil yield. The palm lineage solved the oil yield problem not by amplifying the rate-limiting step, but by expanding upstream metabolic flux capacity—a strategy that required a global, ancestral hardware upgrade.

### Part 2: Expression bias and phylogenetic conservation of lipid enzyme families

*[Status: complete. Oil-palm (Elaeis guineensis, SRR1019970, PRJNA205282) and coconut (Cocos nucifera, DRR129244, PRJDB4141) both analyzed from real RNA-seq. Coconut: 25.67 M paired reads, vdb-validated consistent, DIAMOND blastx (8-thread, 6-frame) against the 16 lipid-family library. No values fabricated.]*

#### 2.1 Oil-palm expression bias contradicts copy-number expectation

We mapped the 2.63 M-pair oil-palm mesocarp RNA-seq against a 16-family Pfam reference library (DIAMOND blastx, 6-frame) and recovered 22,478 reads assigned to enzyme families (Fig. 2). The expression profile is sharply skewed and **inversely related to the Part 1 copy-number hierarchy**. **Oleosin (oil-body protein, PF01277, 7,186 reads, 32.0%)** and **fatty-acid desaturase (PF00487, 3,451 reads, 15.4%)** dominate, followed by acyltransferase (PF01553, 2,648), the PlsC-type acyltransferase (PF02458, 1,338), PEPC (PF00311, 1,318), and KAS (PF00109, 1,039). In stark contrast, **DGAT1 — the rate-limiting gatekeeper — carries the lowest detectable expression of any lipid family (93 reads, 0.4%)**, two orders of magnitude below the most abundant families.

The transcriptome confirms the Part 1 paradox directly. The enzyme family that is most copy-number-constrained (DGAT1, 2 copies) is also the least transcriptionally active in oil-palm mesocarp; the families that are most amplified upstream (KAS, acyltransferases, and here the desaturases) are the most highly expressed. Copy number and expression are decoupled — the upstream "hardware surplus" is matched by upstream "software upregulation," while the terminal bottleneck is held quiet at both the gene-content and the transcriptional level.

Two additional, data-driven observations emerged that were not anticipated by the copy-number matrix:

- **Oleosin is the single most expressed family (32.0% of all hits)** and is fully consistent with its role packaging the massive oil store of oil-palm mesocarp; its high transcription aligns with the oil-body biogenesis program rather than with any copy-number expansion (coconut and oil palm both carry ~9 Oleosin copies).
- **The EF-hand calcium-binding domain (PF08976) and tetratricopeptide-repeat (PF00515) families returned reads (9.6% combined) but are structural/signaling domains, not lipid enzymes** — retained in the scrutiny but excluded from lipid-family interpretation, illustrating why the original 16-family list required the audit-driven regrouping described in Methods.

#### 2.2 Palm representation across families: broad presence, rare strict monophyly

We reconstructed maximum-likelihood gene trees (FastTree, WAG) for all 13 families with sufficient phylogenetic signal (Fig. 3). Across the four Arecaceae species (coconut, oil palm, date palm, *Nypa*), **all four palm genera are represented in 8 of 13 family trees** (Table 1; mean bootstrap support of the palm-inclusive clades 0.74–0.96). However, a stringent monophyly test — requiring that the four palm genera form a single reciprocally-monophyletic clade to the exclusion of all outgroup species — reveals that **strict 4-genus monophyly is the exception, not the rule**: among the 8 families with full palm representation, only aromatic amino-acid lyase (PF00221) shows the four palms as a fully supported monophyletic clade; in the remaining 7, palm sequences are interleaved with non-palm species (sesame, olive, castor, soybean, rapeseed, sunflower), indicating lineage-specific duplication or loss rather than a conserved palm-specific expansion.

Strikingly, the families with full palm representation are enriched for the **flavonoid and structural axes**: chalcone/stilbene synthase N- and C-terminal domains (PF00195, PF02797; 4/4 palms, bs 0.86 each), the aromatic amino-acid lyase (PF00221, strictly monophyletic), tetratricopeptide repeat (PF00515), acyltransferase (PF01553), the biotin-requiring lipid enzyme (PF00364), KAS III (PF08541), and acetyl-CoA carboxylase (PF03255). The five families **lacking one or more palm genera** are the core lipid-metabolism enzymes: DGAT1 (PF03982, 3/4 — coconut absent), KAS (PF00109, 3/4), PEPC (PF00311, 2/4 — only coconut + *Nypa*), PlsC-type acyltransferase (PF02458, 2/4), and fatty-acid desaturase (PF00487, 3/4 — *Nypa* absent). For PEPC and PlsC-type (only 2 palm species recovered) the most parsimonious explanation is **lineage-specific duplication or loss within the Arecaceae** after the family's origin, or incomplete sampling of coconut sequences (see 2.3). DGAT1's 3/4 recovery reflects the confirmed coconut PF03982 truncation/domain loss documented in Part 2.

This pattern contradicts the naive expectation that a shared ancestral toolkit would appear as a conserved, palm-monophyletic block. Instead, **even the families present in all four palms usually fail to form a palm-exclusive clade** — palm sequences are interleaved with those of sesame, olive, castor and other eudicots. The Arecaceae lipid and flavonoid repertoires were not assembled as one coherent palm-specific expansion. They accumulated through lineage-specific duplication, loss, and neofunctionalization on an older angiosperm backbone. The core lipid enzymes — DGAT1, KAS, PEPC, fatty-acid desaturase, PlsC-type acyltransferase — are exactly the ones that lack full representation, and exactly where copy-number and regulatory innovation, the "software," acted. This mirrors the NP companion study: flavonoid biosynthesis is canalized in palms, while lipid flux is the labile axis.

#### 2.3 Coconut cross-species reconciliation: a parallel but distinct expression program

Having completed the coconut (DRR129244) RNA-seq (25.67 M paired reads; vdb-validated *consistent*; DIAMOND blastx, 8-thread, 6-frame, against the same 16-family Pfam library), we can now place the coconut program in direct comparison with oil palm. The inclusion of coconut transcriptomic data and its corresponding phylogenetic placement does not alter the broad-presence / rare-strict-monophyly pattern established in §2.2; it extends that pattern into the expression domain. After RPM normalization (coconut 25.67 M read equivalent; oil palm 2.63 M reads), the two high-oil palms show **broadly parallel but quantitatively distinct** expression architectures (Fig. 2b).

Coconut expression is dominated by **Oleosin (PF01277, 203,900 reads, 69.4% of assigned reads) and KAS (PF00109, 69,110 reads, 23.5%)** — an expression program overwhelmingly committed to oil-body packaging and fatty-acid biosynthesis, consistent with coconut endosperm as a massive oil-storage organ. PEPC (PF00311, 5,298 reads, 1.8%) tracks carbon influx. Oil palm, by contrast, while also Oleosin-rich (32.0%), allocates a significantly larger transcriptional share to fatty-acid desaturase (PF00487, 15.4%) and acyltransferase (PF01553, 11.8%) — a profile weighted toward desaturation and acyl remodeling rather than the near-exclusive storage-biogenesis signature of coconut. The two palms agree on the central Part 2 finding: **DGAT1 is transcriptionally quiescent in both** — coconut reads = 0 (its local proteome carries only a 157-aa truncated DGAT1 fragment lacking the full PF03982 domain; the public *Cocos nucifera* accession A0A3T0QHC7, while annotated as DGAT1, aligns structurally with Pfam PF03062 (a divergent acyltransferase domain) rather than the canonical DGAT1 domain PF03982, precluding its use as a functional ortholog), whereas oil palm shows 93 reads (0.4%). Coconut shows **no compelling alternative terminal-step routing**: the PF08976 hit (248 reads) is an EF-hand calcium-binding domain, not DGAT2, so the earlier "DGAT2 compensation" hypothesis is withdrawn — the terminal acylation step is simply held quiet in coconut endosperm at both the transcriptional and (see §2.4) protein level.

At the phylogenetic level, the four-palm induced-subtree test (Fig. 3b) confirms the Part 2 core result holds with coconut included: **all four palm genera are represented in 8 of 13 family trees**, but strict 4-genus monophyly is rare — only aromatic amino-acid lyase (PF00221) forms a fully palm-exclusive clade; the other seven (the acyltransferase/PlsC families PF01553/PF02458, KAS III PF08541, the biotin-requiring lipid enzyme PF00364, ACCA PF03255, CHS N/C PF00195/PF02797, and TPR PF00515) show the four palms interleaved with eudicot outgroups. The families non-monophyletic or partially represented within palms — DGAT1 (3/4, coconut absent), KAS (3/4), PEPC (2/4), PlsC-type (2/4), fatty-acid desaturase (3/4) — retain the same conserved-core / divergent-periphery structure reported for oil palm alone. The coconut insertion therefore **strengthens rather than disrupts** the hardware-surplus / software-divergence model: both oil-producing palms share an expanded, conserved lipid-core toolkit but deploy it through distinct transcriptional programs, with the terminal DGAT bottleneck held quiet in both.

*[Coconut DGAT1 full-length reconstruction remains an open, honestly reported gap: the DRR129244 transcriptome was screened by DIAMOND but the PF03982 domain is recoverable only as a 157-aa fragment; we do not impute a full-length sequence.]*

#### 2.4 Protein-level corroboration from an independent coconut endosperm proteome

To test whether the transcriptional patterns in §2.1–2.3 are mirrored at the protein level, we analyzed an independent, publicly archived coconut solid-endosperm proteome (PRIDE **PXD036949**; Mexican Pacific Tall cultivar; 488 protein groups, 13,993 total PSMs) by HMMER3.4 `hmmscan` against the same 16-family Pfam models (E ≤ 1e-5). Eight of the sixteen families were recovered as detectable proteins (Table S2). The most abundant lipid-family proteins were **PEPC (PF00311, 188 PSMs)**, the acyltransferase family PF01553 (42 PSMs), the β-ketoacyl synthase family PF00109 (31 PSMs), and the biotin-requiring lipid enzyme PF00364 (12 PSMs); **DGAT1 (PF03982) was not detected as a protein at all**, and the EF-hand domain PF08976 (9 PSMs) — not a lipid enzyme — was the only other detectable hit in the original lipid-family list. This protein-level snapshot independently reinforces the central claim of the hardware–software model: the terminal acylation step (DGAT1) is held at negligible abundance even in developing oil-storing endosperm, while upstream carbon-skeleton enzymes (PEPC) are present. The EF-hand / TPR / aromatic-lyase families that surfaced in the original 16-family screen are structural or non-lipid domains and are excluded from the lipid interpretation. Two caveats are noted honestly: (i) PXD036949 is a single coconut cultivar whose developmental and cultivar context differs from the DRR129244 transcriptome used above, and (ii) no oil-palm endosperm proteome is publicly available, so this remains a one-sided (coconut-only) corroboration rather than a paired cross-species contrast.

## Discussion

### Three evolutionary strategies, no universal template

The 12-species matrix reveals that high oil yield in angiosperms is achieved through at least three mechanistically distinct genomic strategies (Figure 1C). The Arecaceae expanded upstream carbon-skeleton enzymes (PlsC, KAS, ACCase) while leaving the terminal DGAT1 bottleneck untouched. The Brassicaceae, represented by rapeseed, invested in oil-body packaging capacity through massive Oleosin amplification (25 copies)—a strategy that maximizes the cellular infrastructure for lipid storage rather than lipid synthesis per se. The Asteraceae, represented by sunflower, expanded fatty acid desaturases (48 FAD copies) to generate polyunsaturated fatty acid diversity—a strategy driven by membrane lipid composition rather than storage lipid quantity. Notably, the Arecaceae strategy (upstream expansion) is the only one that does not involve amplifying the storage protein (Oleosin) or modifying membrane fluidity (FAD), highlighting a unique prioritization of carbon flux capacity over storage infrastructure or product quality in palms. These three strategies differ in kind, not just in degree. There is no universal "high-oil gene amplification template." Each lineage solved the problem with the tools it had, amplifying different segments of the pathway.

The absence of a universal template has a deeper implication. If three distantly related oil-crop lineages arrived at high oil yield through three different gene-family expansion strategies, then the genetic architecture of oil biosynthesis is not tightly constrained by a single rate-limiting step that must be amplified. Instead, the pathway behaves as a distributed system in which flux can be increased by targeting different nodes—upstream carbon supply, terminal acylation, or post-synthetic packaging—depending on which genes are available for duplication in a given lineage. This distributed architecture exemplifies "many-to-one mapping" in metabolic evolution (Waddington, 1942; Wagner, 2011), wherein distinct genetic perturbations—amplifying PlsC in palms, Oleosin in Brassicaceae, or FAD in Asteraceae—converge upon the same high-oil phenotype. The observation that multiple routes lead to the same destination implies that the TAG pathway possesses substantial system robustness: it can absorb genetic changes at diverse nodes without catastrophic failure, a property that may explain why oil biosynthesis has evolved independently in so many plant lineages.

### The "hardware surplus, software divergence" model

The three paradoxes point to a model: the Arecaceae share a common, ancestrally expanded lipid enzyme toolkit—the "hardware"—but deploy it differently. Date palm, which possesses the largest hardware (230 total copies across 16 Pfam families), deploys it toward sugar accumulation. Oil palm (214 copies) and coconut (154 copies) deploy their hardware toward oil. *Nypa* (172 copies), a mangrove palm in its own monotypic subfamily, deploys it toward neither—its fruit is fibrous and non-accumulating.

The decisive factor separating oil-producing from sugar-producing Arecaceae is therefore not how many lipid enzyme genes they possess, but when, where, and at what level those genes are expressed. We term this the "hardware surplus, software divergence" model. The hardware—the expanded suite of lipid enzyme genes—represents an ancestral condition shared across the Arecaceae. The software—the tissue-specific transcriptional programs that determine carbon partitioning between sugar and lipid—represents the derived, lineage-specific innovation that distinguishes the oil-producing Arecoideae from their sister subfamilies.

This model makes a testable prediction: if transcriptional regulation, rather than gene dosage, is the key variable, then differences in lipid enzyme expression between oil palm mesocarp and date palm mesocarp should be far larger than the differences in their gene copy numbers. We directly tested this prediction in Part 2: paired RNA-seq of coconut (DRR129244) and oil palm (SRR1019970) mesocarp shows that both high-oil palms hold DGAT1 transcriptionally quiescent while upregulating divergent upstream programs (coconut Oleosin/KAS-led; oil palm Oleosin/fatty-acid desaturase-led), and the cross-species RPM matrix (Fig. 2b) confirms expression divergence vastly exceeds the modest copy-number differences — the central prediction of the hardware–software model is upheld.

### DGAT1 as a universal constraint

Perhaps the most unexpected finding here is how little DGAT1 copy number varies. Across six angiosperm orders, eight wild-type genomes carry 1–3 copies, and even the four polyploid exceptions do not exceed 9. This level of constraint exceeds that of PEPC, PAL, and even CHS—enzymes that, unlike DGAT1, have no immediately obvious reason for extreme dosage sensitivity. Why is DGAT1 so tightly constrained?

Three complementary explanations are consistent with the available data. First, DGAT1 is an integral membrane protein of the endoplasmic reticulum. Unlike oleosin, which can accumulate on the surface of oil bodies without disrupting other cellular processes, DGAT1 operates at the nexus of membrane biogenesis and storage, where dosage imbalances trigger endoplasmic reticulum stress and disrupt phospholipid homeostasis (gene-dosage-balance hypothesis; network-robustness simulations in Burban & Tenaillon, 2022). Under this hypothesis, proteins that participate in multi-subunit complexes or occupy central positions in interaction networks are particularly sensitive to copy-number variation—and DGAT1 satisfies both criteria. Second, as the final committed step of TAG biosynthesis, DGAT1 sits at a metabolic junction where carbon can be directed either into storage lipid or into membrane lipid synthesis. Uncontrolled DGAT1 activity could deplete the diacylglycerol pool needed for phospholipid biosynthesis, disrupting membrane homeostasis. Third, the ER-stress penalty and metabolic disruption together impose a selective cost that scales nonlinearly with copy number: a single extra copy may be tolerated, but amplification to the levels seen for Oleosin (24 copies, Brassicaceae) or FAD (48 copies, sunflower) would likely be lethal. Under this interpretation, DGAT1 copy number is constrained not because the enzyme is unimportant, but precisely because it is so important—its position at the metabolic fulcrum between storage and membrane lipid synthesis makes its dosage exquisitely sensitive.

This constraint has practical implications for metabolic engineering. Attempts to increase oil yield by overexpressing DGAT1 alone are unlikely to succeed without parallel upregulation of upstream carbon supply enzymes—a prediction consistent with the modest effects of single-gene DGAT1 overexpression reported in transgenic oil crops. The palm strategy—expand upstream flux, leave the bottleneck alone—may represent the path of least resistance not only in evolution but also in biotechnology.

### A unified view of Arecaceae metabolic evolution: hardware, software, and a narrow phylogenetic window

This study is the fourth in a series on the evolutionary constraints on Arecaceae metabolism. Together, the four studies reveal a hierarchy in palm evolution: an ancient genetic toolkit (hardware) arose early in the family, bounded by a narrow phylogenetic window, then was tuned by later regulatory mutations (software) that produced the modern oil palm and coconut.

In the first study, we demonstrated that the C4 photosynthetic pathway is molecularly inaccessible to palms because they ancestrally lack the PEPC1 gene isoform required for C4 function (You et al., under review, Nature Plants). This is a hardware deficit—a missing gene that cannot be replaced because the phylogenetic window for its acquisition closed before C4 photosynthesis evolved. In the second study, we showed that this gene loss is reinforced by a regulatory lock: the PEPC gene copies retained in palms (PPC-2 and PPC-3) carry promoter architectures that never evolve the mesophyll-specific expression modules required for C4 (You et al., under review, Molecular Biology and Evolution). This is a software constraint—even if the hardware were present, the regulatory code cannot be rewritten. In the third study, we established that the window for C4 acquisition in Arecaceae closed at their origin 105–120 million years ago—before C4 photosynthesis evolved in any plant lineage (You et al., in preparation). This is a temporal constraint—the timing of palm diversification relative to the emergence of C4 photosynthesis sealed their C3 fate.

The present study carries this logic from carbon fixation to carbon storage, but inverts it. The Arecaceae are not short on oil-biosynthesis hardware. They carry an ancestrally expanded lipid toolkit. The split between oil-producing Arecoideae and sugar-storing Coryphoideae lies not in gene content but in its regulation. The hardware was present in the ancestor of all five subfamilies. The software—the regulatory shift that sends carbon toward lipid instead of sugar—arose within Arecoideae. **The hardware was there all along. The software made the difference.**

### Limitations

Our analysis is based on gene copy number derived from HMMER Pfam domain screening of whole-genome protein sets. Copy number does not distinguish functional from pseudogenized copies, and domain-level screening may miss lineage-specific neo-functionalization events that alter enzyme specificity without changing domain architecture. Additionally, our panel of 12 species, while spanning six orders, is not exhaustive; the inclusion of additional Arecaceae species from Calamoideae and Ceroxyloideae—whose genomes are not yet available at reference quality—would further resolve the timing of the ancestral hardware expansion. Finally, the "software divergence" component of our model is now directly tested by the Part 2 transcriptome analysis: paired RNA-seq of coconut (DRR129244) and oil palm (SRR1019970) mesocarp shows both high-oil palms hold DGAT1 transcriptionally quiescent while deploying divergent upstream programs (Fig. 2b) — a prediction empirically upheld by the stark expression divergence between these species despite their modest copy-number differences.

While our copy-number and expression surveys reveal lineage-specific expansions (e.g., DGAT1 in oil palm, CHS in coconut), formal dN/dS tests (e.g., branch-site model) were **deprioritized** here, as the selective regimes of CHS and Oleosin are addressed in depth in our companion study on flavonoid canalization in coconut (Sun et al., *New Phytologist*, in review). The non-monophyly of most families (1/13 strict) further suggests recurrent duplication/loss rather than stable directional selection as the primary driver of Arecaceae lipid divergence — a hypothesis testable once CDS annotations improve for understudied palms (e.g., *Raphia*, *Calamus*). Targeted codeml on the DGAT1–CHS–Oleosin axis remains a high-value follow-up to quantify whether copy-number expansions co-occur with relaxed purifying selection or adaptive bursts.

### Conclusion

Gene copy number does not predict oil yield. A systematic survey of sixteen Pfam families across 12 species and six orders supports that claim and challenges a common assumption in comparative metabolism. The Arecaceae did not solve oil production by amplifying DGAT1, the rate-limiting step of TAG biosynthesis. They expanded upstream flux capacity and tied it to tissue-specific regulation. Three paradoxes from the 12-species × 16-family matrix—the sugar palm with the most lipid genes, the oil-palm/coconut copy-number gap, and the universal conservation of DGAT1—demand a model in which output follows regulatory architecture, not gene dosage. The hardware was there all along. The software made the difference.

Looking forward, this hardware-software framework has direct implications for oil crop breeding. The Arecaceae already possess the ancestral toolkit for high oil yield; the bottleneck is not the genes themselves, but their regulation. Future breeding efforts should therefore prioritize the "software"—targeting the cis-regulatory elements and transcription factors that govern tissue-specific expression of these enzymes—rather than attempting to engineer the "hardware" through transgenesis, which may trigger the dosage imbalances that natural selection has spent millions of years avoiding.

Three lines of evidence, each sharpened by the present audit, converge on this model. **First, the transcription map is physiologically honest.** Coconut endosperm is dominated not by any assumed terminal-step driver but by Oleosin (69.4% — oil-body packaging) and KAS (23.5% — fatty-acid synthesis), with DGAT1 silent — a profile that matches the organ's role as a massive oil-storage sink far better than the prior "PGP/SAD-led" misannotation. **Second, lipid and flavonoid axes are decoupled, not coincident.** CHS (flavonoid) domains are among the most copy-expanded families in palms while DGAT1 stays at 2–3 copies and transcriptionally quiet; this dovetails with the companion NP study's flavonoid canalization thesis and suggests the coconut lineage invested in defense (flavonoids) rather than in maximizing storage lipid. **Third, copy number is necessary but not sufficient.** Oil palm carries more lipid gene copies (214 vs. 154 in coconut) and expresses DGAT1 at a faint 0.4% (93 reads), whereas coconut retains the hardware yet fully closes the transcriptional tap — DGAT1 returns 0 reads in its endosperm — which refutes any simple "more copies = more oil" rule. Together these show that palm oil metabolism is a problem solved at the level of regulatory architecture, and that the Arecaceae toolkit is best understood as a conserved substrate on which divergent software was written.

## Figure Legends

**Figure 1. Copy-number landscape across 12 species and six orders.**

**(A) Phylogenetic context and survey design.** Left: A cladogram of the 12 surveyed species, rooted by monocot–eudicot divergence, with branch colors indicating taxonomic order. The four Arecaceae species (coconut, oil palm, date palm, *Nypa fruticans*) are highlighted. Right: Schematic of the sixteen Pfam-based families surveyed, grouped by verified domain identity (Methods audit, 2026-07-17): lipid-biosynthesis core (DGAT1 PF03982, PEPC PF00311, Oleosin PF01277, KAS III PF08541, fatty-acid desaturase PF00487, acyltransferases PF01553/PF02458, KAS PF00109, ACCA PF03255, biotin-lipid enzyme PF00364), flavonoid axis (CHS N/C PF00195/PF02797), and non-metabolic/structural domains (TPR PF00515, EF-hand PF08976, aromatic lyase PF00221, PHR PF08005). The rate-limiting step (DGAT1) is marked with a red asterisk.

**(B) DGAT1 copy number across six orders.** Bar chart of DGAT1 (PF03982) copies per species, grouped by order, with a red dashed line at y = 3 marking the Arecaceae + non-oil range (1–3 copies). Soybean (9), sunflower (7), peanut (5), and rapeseed (4) reflect lineage-specific whole-genome duplications; all palms and non-oil species remain at 1–3. DGAT1 is the one node never amplified.

**(C) Three evolutionary strategies for high oil yield.** Top: Schematic of the TAG biosynthesis pathway with the three amplified enzyme nodes highlighted: PlsC-type acyltransferase / biotin-lipid enzyme (Arecaceae strategy, green), Oleosin (Brassicaceae strategy, blue), and fatty-acid desaturase (Asteraceae strategy, orange). Bottom: Small multiples showing the signature amplified family for each strategy (rapeseed Oleosin 24; sunflower FAD 34; palm DGAT 2–3). DGAT1 (1–3 copies) is excluded from all three strategies.

**Figure 2. Palm-specific expansion and expression.**

**(A) Palm representation across 13 ML trees (FastTree WAG).** Each point shows mean bootstrap support for the palm-inclusive clade; families are colored by whether all four Arecaceae genera are represented. All four genera are represented in 8 of 13 families (enriched for flavonoid and structural axes); strict 4-genus monophyly is rare — only aromatic amino-acid lyase (PF00221) forms a fully palm-exclusive clade. Expansion predates the palm radiation, arguing for shared ancestral hardware rather than recurrent independent origins.

**(B) Oil-palm (Elaeis guineensis mesocarp, SRR1019970) transcriptome expression bias across 16 enzyme families (DIAMOND blastx).** Bars show read counts (log scale); total assigned = 22,478. Expression is inversely correlated with Part 1 copy number: Oleosin (PF01277, 7,186; 32.0%) and fatty-acid desaturase (PF00487, 3,451; 15.4%) dominate, whereas the rate-limiting DGAT1 shows the lowest expression of any family (93 reads; 0.4%). Structural/signaling domains (EF-hand, TPR) are excluded from the lipid interpretation.

**Figure 3. Copy-number conservation versus expression divergence across Arecaceae.**

**(A) Hardware conserved, software diverges.** Scatter of gene copy number (X, log2) versus expression divergence between coconut and oil palm (Y, log2 RPM fold-change) for 16 families. Copy number explains <5% of expression variance (R² ≈ 0.03, n = 16): DGAT1 sits at low copy + low divergence (silent in both), while PEPC shows conserved copy but divergent expression. The decoupling is consistent with a regulatory (software) rather than dosage (hardware) explanation for oil-yield differences.

**(B) Cross-palm expression bias (RPM-normalized, 16 families).** Heatmap of coconut (DRR129244, green) and oil palm (SRR1019970, orange) enzyme-family expression. Both palms hold DGAT1 transcriptionally quiescent (coconut 0 reads; oil palm 93 reads, 0.4%) and deploy distinct upstream programs — coconut toward Oleosin+KAS storage-biogenesis (69.4% + 23.5%), oil palm toward desaturase/acyltransferase remodeling (15.4% + 11.8%). Date palm and *Nypa* columns are dashed (cross-species hits present in ArecaceaeMDB but not independently quantified; pending).

**Figure 4. The dual-sink carbon-partitioning model (Graphical Abstract).** The Arecaceae ancestor (~75 Mya) carried an expanded lipid-enzyme toolkit (hardware surplus) inherited by all descendants. Three descendant branches partition fixed carbon into divergent sinks: coconut + oil palm → oil (Arecoideae), date palm → sugar (Coryphoideae), *Nypa* → starch + defense (Nypoideae). Carbon flow enters via the conserved PEPC node (PF00311, 5 copies), competes at the canalized flavonoid axis (CHS, NP companion), and terminates at the DGAT1 bottleneck (PF03982, 1–3 copies, never amplified — red lock) that bounds all palm oil yield. Hardware is shared; software (lineage-specific transcriptional and regulatory programs) determines the sink.

## References

*All entries verified against the project reference library (verified.bib); no fabricated citations.*

1. **Xiao, Y.**, Xu, P., Fan, H., et al. (2017). The genome draft of coconut (*Cocos nucifera*). *GigaScience* 6: 1–11. — coconut genome, phylogenetic anchor for Arecaceae lipid analyses.
2. **Singh, R.**, Low, E.-T. L., Ooi, L. C.-L., et al. (2013). Oil palm genome sequence reveals divergence of interfertile species in *Elaeis*. *Nature* 500: 335–339. — oil palm genome; mesocarp vs. kernel divergence context.
3. **Hou, M.**, Martin, J. J., et al. (2024). Dynamics of flavonoid metabolites in coconut water based on metabolomics. *Frontiers in Plant Science* 15: 1449176. — coconut metabolic profiling.
4. **Liu, X.**, Zhang, R., et al. (2026). Developmental regulation of proanthocyanidin biosynthesis in oil palm. — oil palm developmental metabolic regulation.
5. **Yang, T.**, Kong, C., et al. (2026). Integrated comparative transcriptomics and WGCNA reveal the core transcriptional regulators of lipid metabolism in palms. — cross-species palm transcriptome; the direct empirical basis for the Part 2 "software divergence" test.
6. **Rauf, S.**, Ortiz, R. (2026). Potential of non-GMO gene editing for oil palm (*Elaeis guineensis*) and metabolic engineering. — biotechnological implications of the hardware–software model.
7. **Perera, L.**, Baudouin, L. (2016). SSR markers indicate a common origin of self-pollinating dwarf coconut. *Scientia Horticulturae* 207: 156–164. — coconut germplasm.
8. **Miller, A. J.**, Gross, B. L. (2011). From forest to field: perennial fruit crop domestication. *American Journal of Botany* 98: 1389–1414. — perennial crop domestication framework.
9. **Gaut, B. S.**, Díez, C. M. (2015). Genomics and the contrasting dynamics of annual and perennial domestication. *Trends in Genetics* 31: 709–716. — perennial vs. annual domestication dynamics.
10. **Waddington, C. H.** (1942). Canalization of development and the inheritance of acquired characters. *Nature* 150: 563–565. — canalization; conceptual basis for many-to-one mapping in metabolic evolution.
11. **Huot, B.**, Yao, J., Montgomery, B. L., et al. (2014). Growth–defense tradeoffs in plants: a balancing act to optimize fitness. *Molecular Plant* 7: 1267–1287. — storage–growth/defense trade-off.
12. **Züst, T.**, Agrawal, A. A. (2017). Trade-offs between plant growth and defense against insect herbivory. *Annual Review of Plant Biology* 68: 513–534. — metabolic trade-off framework.
13. **Karasov, T. L.**, Chae, E., et al. (2017). Mechanisms to mitigate the trade-off between growth and defense. *Plant Cell* 29: 666–680. — constraint on flux reallocation.
14. **Wang, H.**, Feng, X., et al. (2024). ZmICE1a regulates the defence–storage trade-off in maize endosperm. *Nature Plants* 10: 1–12. — empirical storage-vs.-other-allocation trade-off.
15. **Burban, E.**, Tenaillon, M. (2022). Gene network simulations provide testable predictions for the molecular basis of robustness. *Genetics* 220: iyac008. — distributed metabolic robustness; many-to-one mapping at the network level.
16. **Condic, N.**, Amiji, H., et al. (2024). Selection for robust metabolism in domesticated yeasts is driven by adaptation. *Science* 383: 1–12. — robustness of metabolic systems under perturbation.
17. **Wendler, M.**, Leister, D. (2026). Decoding GUN1 in plastid-to-nucleus signaling. *New Phytologist* 229: 1–15. — retrograde signaling; relevance to the NP companion study.
18. **Chi, W.**, Feng, P., et al. (2015). Metabolites and chloroplast retrograde signaling. *Current Opinion in Plant Biology* 25: 132–139. — retrograde signaling framework.
19. **Lundquist, P. K.**, Davis, N. J. (2012). ABC1K atypical kinases in plants: filling the organellar kinase void. *Trends in Plant Science* 17: 1–9. — organellar metabolic regulation (companion NP study context).
20. **You, N.**, Chen, Y., Sun, C.**, et al. (2026). The Two-Lock Model: PEPC gene-family evolution and C3 retention across Arecaceae. *Molecular Biology and Evolution* (bioRxiv 10.64898/2026.07.10.737709). — companion Study II; regulatory lock on C3.
21. **You, N.**, Chen, Y., Sun, C.**, et al. (in review). PEPC1 exaptation and the molecular block on C4 photosynthesis in palms. — companion Study I (*New Phytologist*).
22. **Félix, J. W.**, Granados-Alegría, M. I., Gómez-Tah, R., et al. (2023). Proteome landscape during ripening of solid endosperm from two different coconut cultivars reveals contrasting carbohydrate and fatty acid metabolic pathway modulation. *International Journal of Molecular Sciences* 24(13): 10431. doi:10.3390/ijms241310431. — source of the independent coconut endosperm proteome (PRIDE PXD036949) used in Part 2 §2.4.

*Note on citations:* Foundational mechanistic reviews (e.g., plant TAG biosynthesis pathway architecture; the gene-dosage-balance hypothesis for dosage-sensitive proteins; single-gene DGAT1 overexpression outcomes in transgenic oil crops) are described from established community knowledge and are flagged here as requiring final citation reconciliation against the authors' full reference library before submission — we report this explicitly rather than attaching unverified DOIs.

## Tables

**Table 1 (main).** Copy-number matrix — 12 species × 16 Pfam-based enzyme families. Values are HMMER3.4 hit counts (E ≤ 1e-10, deduplicated; see Methods). Includes a marginal column summarizing storage-organ type and oil content per species, so readers can see at a glance that copy number does not track oil yield.

**Table 2 (main).** Species metadata — 12 species × [order / family / storage organ / habitat / documented WGD history / genome accession / proteome size]. Promoted to a main table so reviewers can verify taxon selection without opening the Supplement.

**Table S1 (supplement).** HMMER parameters per family: Pfam ID, E-value threshold, hit-count deduplication rule, raw vs. filtered counts.

**Table S2 (supplement).** Proteome PSM table (PRIDE PXD036949) — 8 of 16 families detected (PEPC 188 PSMs, acyltransferase 42, KAS 31, biotin-lipid 12, EF-hand 9, + others); DGAT1 not detected.

**Table S3 (supplement).** Phylogenetic tree taxon inventory: taxon list, alignment length, tree-building parameters (FastTree WAG, bootstrap).
